<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Inside Elastic - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Inside Elastic - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/search-labs/blog/category/inside-elastic</link>
    </image>
    <link>https://www.elastic.co/search-labs/blog/category/inside-elastic</link>
    <atom:link href="https://www.elastic.co/search-labs/rss/category/inside-elastic.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Tue, 29 Sep 2026 14:29:06 GMT</lastBuildDate>
  <item>
    <title><![CDATA[GPU-accelerated vector indexing in Elasticsearch with NVIDIA cuVS: 138M vectors in under 10 minutes]]></title>
    <description><![CDATA[Moving index builds to the GPU leaves the CPU free for queries, which is how vector indexing throughput went up 7x and p90 search latency fell 6x while indexing ran, with no change to recall.]]></description>
    <content:encoded><![CDATA[<p>Modern enterprise applications are ingesting terabyte- to petabyte-scale unstructured data to power semantic search, large language model–based (LLM-based) retrieval augmented generation (RAG), and recommender systems. At this scale, <a href="https://www.vastdata.com/blog/powering-enterprise-ai-high-velocity-vector-search-sql">vector indexing on CPUs can take days or even weeks</a>, slowing experimentation and making large index updates operationally expensive. This is especially challenging for workloads that require regular index rebuilds, such as those driven by frequent data updates or frequent embedding model updates. For example, several ecommerce teams refresh their product catalog nightly and train their own embedding models, each requiring an index rebuild.</p><p></p><p>CPU-based indexing can also degrade search performance when indexing and search happen simultaneously, especially in online systems, like ecommerce platforms and ad-serving pipelines. This happens because index builds consume significant CPU resources, leaving fewer cycles available for queries and driving up search latency and latency variance. Even after indexing completes, latency can remain high due to index fragmentation. Customers often run a forced merge to consolidate segments, but it can be slow on CPUs, further delaying recovery to the expected search latency by hours to days.</p><p></p><p><a href="https://www.elastic.co/elasticsearch/vector-database">Elasticsearch</a> has <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/gpu-vector-indexing">introduced GPU-accelerated vector indexing</a> powered by <a href="https://developer.nvidia.com/cuvs">NVIDIA cuVS</a>. In this blog, we show how Elasticsearch achieves up to 7x faster vector indexing throughput by offloading hierarchical navigable small world (HNSW) index construction to GPUs, indexing 138 million vectors in under 10 minutes using eight NVIDIA RTX PRO 6000 GPUs.</p><p></p><h2>Indexing over 100M Vectors in under 10 minutes</h2><p>Figure 1 shows vector indexing throughput on 138 million 1,024-dimensional vectors from the <a href="https://github.com/elastic/rally-tracks/tree/master/msmarco-v2-vector">MS MARCO</a> dataset, representing roughly 4 TB of multimodal documents (Figure 1). The benchmark was run with <a href="https://github.com/elastic/rally">Rally</a>, Elasticsearch’s macro-benchmarking framework. The test ran on a server with 8 NVIDIA RTX Pro 6000 GPUs with a 2 socket AMD EPYC 9555. GPUs were turned on for the GPU test and turned off for the CPU test. The result was that Elasticsearch achieved 7x higher indexing throughput, lowering time to index 138M vectors from 1 hour on CPUs to under 10 minutes on GPUs.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt05167c135e5e0727/6ab4e86f7a4851d33df9803f/vector-indexing-throughput.png" alt="Bar chart of vector indexing throughput: 281K docs/s on GPU versus 38K docs/s on CPU for 138M vectors" /><p></p><h3>Does GPU vector indexing change recall?</h3><p>Index acceleration isn’t helpful if indexes built on GPU have lower search quality compared to CPUs. The throughput-versus-recall curves for GPU-built and CPU-built indexes overlap, showing that GPU indexing delivers the same recall profile as Elasticsearch’s CPU-based indexing path, with minimal accuracy or search-time performance penalty (Figure 2).</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdbb6f3deb0bcaec9/6ab4e8ebafd558dce24a0a33/qps-vs-recall.png" alt="Chart showing QPS versus recall curves overlapping for GPU-built and CPU-built vector indexes at k=100" /><p></p><h3>Search latency under concurrent indexing load</h3><p>Offloading indexing to the GPU improved search latency on the CPU by 6x when running indexing and search simultaneously (Figure 3) because indexing on GPUs frees CPU resources for faster search latency. This allows a single Elasticsearch cluster to support real-time ingestion and low-latency retrieval at the same time, without forcing a tradeoff between freshness and search performance.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd34d2464499b03e4/6ab4e9637c5f023f526f3c67/p90-search-latency.png" alt="Bar chart of p90 search latency during indexing: 7 ms with a GPU-built index versus 45 ms on CPU" /><p></p><h3>Force merge time for HNSW indexes on GPU versus CPU</h3><p>Additionally, GPU-accelerated indexing in Elasticsearch reduced forced merge time from roughly four hours to five minutes, helping restore low-latency search in real-time (Figure 4). Note: Force merge was used to merge each shard into four segments.</p><p></p><p>Here are the full results for GPU versus CPU vector indexing in Elasticsearch:</p><p></p><p><strong>Metric</strong></p><p><strong>CPU</strong></p><p><strong>GPU</strong></p><p><strong>Change</strong></p><p>Index build time, 138 million vectors</p><p>~1 hour</p><p>Under 10 min</p><p>7x throughput</p><p>p90 search latency under indexing load</p><p>Baseline</p><p>6x lower</p><p>6x</p><p>Force merge, four segments per shard</p><p>~4 hours</p><p>~5 min</p><p>~48x</p><p>Recall at target</p><p>95%</p><p>95%</p><p>Unchanged</p><p></p><p></p><h2>What GPU vector indexing means for Elasticsearch users</h2><p>Elasticsearch’s integration of NVIDIA cuVS brings GPU-accelerated vector indexing to production AI search with over 84% reduction in search latency to deliver faster ingest with a 640% increase in indexing throughput with a lower CPU overhead and limited infrastructure tradeoffs. This unlocks a new class of applications, such as faster RAG pipelines, semantic search, multimodal retrieval, and fraud detection at scale.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt54b3129de4981654/6ab4e9c7bf5efbe04ad38b76/force-merge-time.png" alt="Bar chart of force merge time for a 138M vector index: 5 minutes on GPU versus 249 minutes on CPU" /><p></p><h2>Get started with GPU vector indexing in Elasticsearch</h2><p>Install Elasticsearch with GPU-accelerated vector indexing by following <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/gpu-vector-indexing">these instructions</a>. Replicate the above benchmarks by running<a href="https://github.com/elastic/rally-tracks/tree/master/msmarco-v2-vector"> rally msmarco track</a> with parameters in <a href="https://github.com/elastic/rally-tracks/issues/1188">the GitHub issue</a>. Learn more about GPU-accelerated vector indexing, search, and preprocessing by visiting<a href="https://github.com/NVIDIA/cuvs"> NVIDIA cuVS</a>. </p><h2>Acknowledgments</h2><p>The authors would like to thank Chris Hegarty, Gilad Gal, and Mayya Sharipova from Elastic, as well as Nathan Stephens from NVIDIA, for their contributions to this article.</p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/vector-indexing-gpu-elasticsearch-cuvs</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/vector-indexing-gpu-elasticsearch-cuvs</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Index Data]]></category>
    <dc:creator><![CDATA[Bao Tong,Manas Singh,Corey Nolet]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2415fbdb68d01a49/6ab4e2a39a8267823dc47d59/NVIDIA.jpg" length="0" type="image/jpeg"/>
    <pubDate>Fri, 25 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[One button, three places: How we rebuilt Kibana's page headers with stricter APIs]]></title>
    <description><![CDATA[We gave Kibana's shared shell typed contracts, which is how design system governance became the default, and why the new page headers have no breadcrumbs.]]></description>
    <content:encoded><![CDATA[<p>We replaced the open-ended React APIs behind every page header in Kibana with typed contracts, which put design system governance in the shell itself. Dozens of teams no longer have to get each header right by hand. Now a page declares what a control means, and the shared shell decides how it looks and where it goes. We removed breadcrumbs along the way, because the redesigned navigation already shows you where you are. The new chrome is live in Elastic Cloud Serverless today and ships to Elastic Cloud Hosted and self-managed in 9.6.</p><h2>What is the Kibana chrome?</h2><p>The chrome is everything that frames an application in <a href="https://www.elastic.co/kibana">Kibana</a>: the global header and the navigation, along with the header of the page that you’re currently on. Every application renders inside it. Because the chrome is shared, its APIs decide how consistent the whole product feels and how hard the next redesign will be.</p><h2>What inconsistent page headers looked like to users</h2><p>The clearest way to see the problem is to look at what users saw:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt63a6acf0d30ae807/6aaa27addcc13f41b2527962/unnamed_(1).png" alt="Three Kibana pages with different page header layouts, showing inconsistent action button and link placement" /><p></p><p>Take something as basic as a primary action, the one button a page most wants you to click, like <strong>Create index</strong>. Before this work, primary actions lived in three different places, depending on the page:</p><ul><li><p>In the old app header.</p></li><li><p>In the page template header.</p></li><li><p>Somewhere in the page body itself.</p></li></ul><p>Common links had the same problem, with share, feedback, and documentation links appearing in separate spots with different styling from one app to the next. Breadcrumbs were misconfigured on many pages. For example, some showed a breadcrumb for each tab within a page, and some weren't clickable. Still others repeated the title of the page you were already on.</p><p>Every one of these was the output of a well-intentioned team using a flexible API in the way it’s usually used. Users had to relearn each page's layout, and every attempt to redesign the chrome had to account for hundreds of local variations.</p><h2>Goals for the page header redesign</h2><p>The goal of the redesign was simple; that is, pages should express what they need (for example, a title, badge, or primary action), and the shared shell should decide how it looks and where it goes. Common links and primary actions should get one home, everywhere. If it's the primary action, it's always in the same place, styled the same way. And consistency should stop being something that each of dozens of teams has to get right by hand, forever.</p><p>Getting there wasn't a restyling exercise; it required changing the contract between applications and the platform.</p><h2>How typed props replaced EuiPageHeader's open React nodes</h2><p>Each page header was easy to build in isolation, and the local choices piled up into visible differences across the product: inconsistent spacing, title styling, and action placement from one page to the next.</p><p>The fix, everywhere this showed up, was the same mechanism: stricter APIs. Instead of accepting arbitrary React nodes, a control declares what it means through typed props, and the shared shell decides how it looks and where it goes. That one change buys two things at once: consistency, and governance over what the shared surface is allowed to contain.</p><p>Take the page header. With <code>EuiPageHeader</code>, the API exposed layout settings and accepted React nodes for nearly every visible area:</p>&lt;EuiPageHeader
  pageTitle={
    &lt;&gt;
      Index Management
      &lt;EuiBadge color="accent"&gt;Beta&lt;/EuiBadge&gt;
    &lt;/&gt;
  }
  description={
    &lt;&gt;
      View and manage your Elasticsearch indices.{' '}
      &lt;EuiLink href={docsUrl}&gt;Learn more&lt;/EuiLink&gt;
    &lt;/&gt;
  }
  bottomBorder
  alignItems="top"
  responsive={false}
  tabs={tabs}
  rightSideItems={[
    &lt;RefreshButton onClick={onRefresh} /&gt;,
    &lt;CreateButton onClick={onCreate} /&gt;,
  ]}
/&gt;<p>Because <code>pageTitle</code>, <code>description</code>, and <code>rightSideItems</code> all accept arbitrary React nodes, every team filled them in differently. One page put a badge next to the title, while another styled its own pill. One team's actions were plain buttons in one order; the next team's were a different mix in another order, some collapsing into a menu and others not. Everyone used the same shell component, yet the headers looked and behaved like they came from different products. The shared component guaranteed the wrapper, not what went inside it.</p><p>The <code>AppHeader</code> API closes that off. The same header is expressed as typed props:</p>&lt;AppHeader
  title="Index Management"
  badges={[
    {
      label: 'Beta',
      color: 'accent',
      tooltip: 'This feature is in beta.',
    },
  ]}
  description={{
    text: 'View and manage your Elasticsearch indices.',
    learnMoreUrl: docsUrl,
  }}
  tabs={tabs}
  menu={{
    primaryActionItem: {
      id: 'create',
      label: 'Create index',
      iconType: 'plusInCircle',
      run: onCreate,
    },
    items: [
      {
        id: 'refresh',
        label: 'Refresh',
        iconType: 'refresh',
        run: onRefresh,
      },
    ],
  }}
/&gt;<p>Now there’s only one way to express each part, so every page renders it the same. Badges live in their own typed collection, separate from the title. A description carries its text and, optionally, a "Learn more" URL. Actions declare whether they’re primary or secondary, and the component decides how they look and when they collapse. Plus, richer controls fit the same shape. An editable title or a favorite toggle is a structured config rather than a bespoke React tree:</p>&lt;AppHeader
  title={{
    text: indexName,
    onSave: renameIndex,
  }}
  favorite={{
    status: favoriteStatus,
    onToggle: toggleFavorite,
  }}
/&gt;<p>The split is consistent throughout; the application supplies state and behavior (the current name, what happens on rename, how to toggle a favorite) and the header owns presentation. It renders the controls, handles keyboard interaction, shows validation errors, and reflects pending changes. Because those roles are explicit, layout and responsive behavior stay inside the shared component, and applications don't need a coordinated update every time the header changes.</p><h3>Why a shared UI shell needs a closed set of controls</h3><p>The global shell had a similar problem that was one step more open-ended. Any plugin could register arbitrary content on the left or right and pick its position with a number, and no conversation with the platform team or designers was required:</p>chrome.navControls.registerLeft({
  content: &lt;ProjectPicker /&gt;,
});

chrome.navControls.registerRight({
  order: 10,
  content: &lt;AiAssistantButton /&gt;,
});

chrome.navControls.registerRight({
  order: 20,
  content: &lt;FeedbackButton /&gt;,
});<p>That <code>registerRight</code> API is an open door. Any team can add a control, and the header fills up with elements no one designed together. The controls compete for space and carry their own styling, and their order is decided by whoever picked the larger number. The shell only knows that one React node follows another; it can't tell that one opens an AI assistant and another collects feedback, so it can't reason about them or keep them coherent.</p><p>The new shell replaces that open canvas with a closed set of named roles under <code>chrome.controls</code> and <code>chrome.help</code>. The old registry hasn't gone anywhere yet, because the classic header still runs alongside the new one while applications migrate, but nothing in the new shell is reachable through it:</p>chrome.controls.projectPicker.set(&lt;ProjectPicker /&gt;);
chrome.controls.aiButton.register({ content: &lt;AiAssistantButton /&gt; });
chrome.controls.globalSearch.set({ onClick: openGlobalSearch });
chrome.help.registerFeedbackHandler(openFeedback);<h3>Global chrome header versus application page header</h3><p>Part of what made the old chrome hard to evolve was that "the header" was really several things tangled together. The redesign draws a hard line between two surfaces with different owners:</p><p>
</p><p><strong>Chrome (global) header</strong></p><p><strong>Application header</strong></p><p>Owner</p><p>The platform</p><p>The page</p><p>Scope</p><p>What's true everywhere in Kibana</p><p>What's true on this page</p><p>Contents</p><p>Navigation, project or deployment picker, search, help, AI assistant, feedback</p><p>Title, badges, description, tabs, page actions</p><p>How it's populated</p><p>Named slots under <code>chrome.controls</code> and <code>chrome.help</code></p><p><code>AppHeader</code> typed props, rendered directly or set via <code>chrome.appHeader.set()</code></p><p>Because each surface has one owner and a typed contract between them, the platform can redesign the chrome without auditing hundreds of pages, and an application can evolve its header without colliding with global controls. Setting a header config returns a cleanup callback, so an application tears down its own contribution when it unmounts.</p><h2>Design decisions that strict APIs force</h2><p>Stricter APIs forced design decisions that flexible APIs had let everyone defer. Two are worth calling out.</p><h3>Why we removed breadcrumbs from Kibana's navigation</h3><p>Removing breadcrumbs from the chrome was one of the more debated calls, and the state of the data made it easier than expected. In practice, breadcrumbs across Kibana were widely misconfigured:</p><ul><li><p>Some pages generated a breadcrumb for every tab on the page.</p></li><li><p>Some breadcrumbs weren't clickable.</p></li><li><p>Some repeated the title of the current page.</p></li></ul><p>They added visual weight without reliably adding orientation.</p><p>At the same time, the redesigned navigation already conveys hierarchy. The primary and secondary navigation show you where you are and give you a direct way back up. Keeping breadcrumbs would have meant maintaining a third (and frequently wrong) representation of the same information, so we stopped drawing the trail and let the navigation do that job.</p><p>The breadcrumb API itself didn't disappear. Applications still register breadcrumbs, and the new shell reads them to derive the back button and to fall back to a page title when an app hasn't supplied one. The data changed jobs rather than going away, which is why migrating an app to the new header rarely starts with ripping breadcrumbs out:</p><p><strong>Before:</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt45c047fdfcad7c14/6aaa27d97d925c9a3baf92a0/unnamed.png" alt="Kibana page header with a breadcrumb trail repeating the GenAI Settings page title" /><p><strong>After:</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6b04544a35718f97/6aaa27fe92336926efb1c2d0/unnamed.png" alt="the same Kibana page header with breadcrumbs removed, showing only the page title" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2ecbc17acb523ead/6aaa28219936f5d5f78b976d/unnamed.png" alt="Kibana breadcrumbs repeating a page title and tab name, annotated to show misconfigured navigation hierarchy" /><h3>How many action buttons a page header should show</h3><p>The other recurring debate was priority. With every page's actions now flowing through one menu structure, which buttons deserve to be visible, and in what order? Priorities shift as products evolve, so this conversation is ongoing.</p><p>For launch, we made one deliberate simplification: we locked the number of buttons that can appear in the app menu. A page gets one primary action and a bounded set of secondary actions before the rest collapse into a menu. The constraint keeps any single page from dividing the user's attention across a wall of buttons, and it makes the priority conversation explicit instead of letting it be settled by whoever adds the next button.</p><p>We went through several iterations to land here. Initially, we had many visible actions. A page could have up to three buttons, plus a secondary action to the left of the primary one. That took up too much space, so we gradually simplified the layout by removing the secondary action and reducing the number of visible buttons. We also made the overflow menu more structured. Some items, such as feedback and documentation, now have a fixed place in the footer of the overflow menu.</p><h2>What strict design system governance costs teams</h2><p>Under the old API, a plugin could add a novel UI by choosing a side, picking an order, and mounting a React tree. Under the stricter API, a new kind of control may need a shared capability before a plugin can add it. That takes more work up front and forces a design decision that teams could previously avoid.</p><p>We accepted the cost of stricter APIs because the alternative pushed it into every later redesign. As long as arbitrary React trees mounted at arbitrary points, every change to the shell's layout or accessibility had to account for all of them.</p><p>That’s the trade we made, and it pays back on both fronts. Consistency stops being something each team has to get right by hand; there’s one way to express a control, so pages match by default. And the shared surface stays governed. Its set of controls is a deliberate design decision, not whatever accumulated at the edges. Applications describe what their controls mean, and the shell is free to decide how they look, both today and in the next redesign.</p><h2>Where the redesigned Kibana chrome is available</h2><p>The redesigned chrome is available now in <a href="https://www.elastic.co/cloud/serverless">Elastic Cloud Serverless</a>. Open any project, and you're already using it. For <a href="https://www.elastic.co/cloud">Elastic Cloud Hosted</a> and self-managed users, it ships in 9.6.</p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/design-system-governance-kibana-page-headers</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/design-system-governance-kibana-page-headers</guid>
    <category><![CDATA[Kibana]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Developer Experience]]></category>
    <dc:creator><![CDATA[Anton Dosov,Ryan Keairns,Krzysztof Kowalczyk,Alex Marhaba]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc5228b5443fe9d20/6aaa28eb27e4f0a047befaa4/unnamed.png" length="0" type="image/png"/>
    <pubDate>Wed, 16 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Columnar storage isn't a columnar database. What Columnar mode brings to Elasticsearch]]></title>
    <description><![CDATA[Elasticsearch has stored data in columns since 2013, but adding full columnar database capabilities required a new mode.]]></description>
    <content:encoded><![CDATA[<p>When we wrote that Elasticsearch is becoming a columnar database, the sharpest reply we got was that it already is one. That reply is correct on the facts. Doc values, the per-field column store that Elasticsearch inherited from Lucene, arrived in 2013, and nearly every aggregation, sort, and query in Elasticsearch Query Language (ES|QL) has read them since Elasticsearch 2.0 made them the default. Each field's values sit together in their own file on disk. So the interesting question is what else a columnar database needs (rather than whether we store columns), and the answer turns out to be five things.</p><p>Doc values were built to make aggregations, sorting, and grouping possible on a document engine, and they do that job well. <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-storage">Columnar Mode</a> changes what the columns are for, and this post walks through the five properties that separate storing columns from being a columnar database.</p><h2>Five properties separate a column store from a columnar database</h2><h3>Doc values were an optimization on top of <code>_source</code></h3><p>For most of the last decade, the original JSON document was the source of truth and the columns were a derived convenience. That ordering has consequences throughout the engine.</p><p>Because the engine could always fall back to <code>_source</code>, per-field storage was allowed to be lossy. Text fields had no doc values at all, since text could be reread from the stored document when needed. Even synthetic source, which reconstructs a document from its fields rather than storing a copy, sometimes reads from row-shaped structures to stay faithful to the JSON that arrived, with values that exceed <code>ignore_above</code> and fields that arrived unmapped going into stored fields.</p><p>The result is a clear contract; whatever JSON you send, you get back, and the columns accelerate everything else. For an engine whose job is to return your documents, that’s the right way round. Columnar Mode inverts it. Every field stores itself exactly once as doc values, doc values cannot be turned off, text fields get doc values, too, and the document is reconstructed from the columns when something asks for it.</p><h3>How dictionary encoding handles high-cardinality data</h3><p>Sorted doc values, the default for keyword fields, store a dictionary of distinct values plus one ordinal per document pointing into it. This is an excellent trade when values repeat. A <code>host.name</code> field drawn from a hundred machines, or a status code that’s almost always 200, compresses beautifully and groups quickly.</p><p>It works like the index cards in a warehouse. When 50 crates hold the same product, one card and 50 pointers beats writing the product name 50 times. When 50 crates hold the same product, you can store the product name in the index with the list of 50 crate IDs. When every crate holds something unique, you might as well just put the product name on the crates; the index will help you find what crate you want, but it won't save ink.</p><p>High-cardinality fields describe a lot of real data, including URLs and trace identifiers, along with message bodies. Columnar Mode, which skips the dictionary and compresses the values in blocks instead, uses binary doc values for high-cardinality strings. Which of the two a field gets isn’t something you configure. The engine decides per field, based on the values it sees, so each column is encoded for the data it actually holds instead of one default applied to every field. Pure columnar systems have long carried cardinality in their type system, but usually as something you declare, and you own the consequences when the data shifts underneath it. Here, it’s the engine's job.</p><h3>Why every field builds an inverted index by default</h3><p>By default, a keyword field also builds an inverted index and a numeric field also builds a BKD tree. That happens on every field because at write time the engine doesn’t know which capability you’ll want at read time, significantly increasing the footprint of each field. Those structures also have to be rebuilt during segment merges, which costs CPU exactly when ingest is heaviest.</p><p>Our time series engine (TSDB) is proof of what happens when you stop paying for capability that the workload doesn’t use. Replacing the indices on <code>@timestamp</code> and dimension fields with <em>doc value skippers</em>, which are sparse structures holding the minimum and maximum value for each block of documents, removed 10 bytes of the original 25 bytes per OpenTelemetry (OTel) data point. There was no measurable query regression on time range and dimension filters, and indexing CPU dropped by about 10%, as a bonus.</p><p>Columnar Mode generalizes that default. Fields aren’t indexed unless something needs them to be, with only text-mapped fields keeping their inverted index for fast free-text search.</p><h3>Metadata fields like _id and _routing were row-shaped</h3><p>The fields you never think about followed the same document-first design. The <code>_id</code> field was a stored field plus an inverted index. Custom <code>_routing</code> was a stored field. Sequence numbers were kept for optimistic concurrency control, regardless of whether a workload ever updated a document.</p><p>TSDB deals with all three. It synthesizes <code>_id</code> from the <code>_tsid</code> and <code>@timestamp</code> values that already identify a data point, using a segment-level bloom filter to catch duplicates, which removes 5 bytes per data point with no loss of functionality. It trims sequence numbers once replication no longer needs them, which removes 4 bytes. Add a codec block size increase from 128 to 512 elements for another 2 bytes, and those four changes contribute across versions 9.1 through 9.4 to the 21 bytes that took OTel metrics <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-metrics-engine-30x-faster-prometheus">from 25 bytes per data point down to 3.75</a>.</p><p>Columnar Mode makes those ideas general rather than metrics-specific. By general availability (GA), all metadata fields will store themselves as doc values, while we plan to follow up and add a sort id mode synthesizing the identifier from the index sort fields, in addition to derived fields that will generalize what <code>_tsid</code> does for time series to any set of fields.</p><h3>Columnar query execution in the ES|QL compute engine</h3><p>A column store only pays off if the engine reads it as columns. Aggregations inherited the document-at-a-time shape from search, which is the natural fit for an engine built around documents. Reading columns instead lets the engine hand a whole block of values to a single instruction, and that’s where the numbers below come from.</p><p>The ES|QL compute engine changed that shape, and TSDB again shows the size of the effect:</p><ul><li><p><strong>Vectorized execution</strong> of time series aggregations was worth up to 8x on its own.</p></li><li><p><strong>Decoding on-disk data</strong> straight into the primitive arrays the engine aggregates over, with no intermediate copies, was worth roughly another 10x.</p></li><li><p><strong>Constant blocks</strong> turned repeated values into a form of in-memory run-length encoding.</p></li><li><p><strong>Filter pushdown</strong> moved filters down to Lucene, where skippers can discard whole blocks unopened.</p></li></ul><p>Together with the rest of the block-level query work, query latency improved by up to 160x compared to earlier versions.</p><p>That work continues. Skipper-aware operators, aggregations that group on ordinals and convert to real values as late as possible, and richer per-block summaries are all in progress, and they benefit every index mode because every mode reads doc values underneath.</p><h2>What Columnar Mode changes for logs and analytical data</h2><p>Storing values in columns is a storage detail. A columnar database needs five things:</p><ol><li><p>The columns are the only copy of the data.</p></li><li><p>Each column is encoded for the data it actually holds.</p></li><li><p>Metadata is columnar, too.</p></li><li><p>Fields add indices only when something needs them.</p></li><li><p>The query engine processes blocks of values rather than records or documents.</p></li></ol><p>TSDB reached all five for metrics in Elasticsearch 9.4, which is why the numbers in this post come from metrics rather than from a slide. Columnar Mode applies the same treatment to logs and security telemetry, along with analytical data. It’s in technical preview in Elasticsearch 9.5, with GA targeted for 9.7.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt723063cac8566f76/6aa8fb5a5ceda93f62dae858/unnamed.png" alt="Elasticsearch doc values, inverted index and _source across standard index mode, LogsDB and Columnar Mode" /><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-doc-values-columnar-database</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-doc-values-columnar-database</guid>
    <category><![CDATA[Lucene]]></category>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Yannis Roussos]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt25aa087300badaae/6ab3e7938ab0c04cc32254df/diagram-one-field-three-structures.webp" length="0" type="image/webp"/>
    <pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[How we built PromQL into Elasticsearch]]></title>
    <description><![CDATA[PromQL runs on the same Elasticsearch compute engine as ES|QL, with no plugin and no separate process to operate. Getting there meant changing how the engine evaluates time windows and builds grouping keys.]]></description>
    <content:encoded><![CDATA[<p>More than 80% of the Prometheus Query Language (PromQL) queries in our real-world corpus run on Elasticsearch without modification. Elasticsearch 9.5 makes the PromQL and the Prometheus-compatible API generally available (GA), so you can ingest Prometheus metrics with<a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/tsds-ingest-prometheus-remote-write"> remote write</a> and query them through the<a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api"> Prometheus HTTP APIs</a> or the<a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/promql"> PROMQL</a> command in Elasticsearch Query Language (ES|QL).</p><p>PromQL compiles to the same compute engine that runs ES|QL and inherits its planner and distributed execution, along with its release process. We didn’t build a second engine for this, and there’s no plugin to install.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9e310070ec244c5d/6aa3928e224d356c5e0dd18e/1.png" alt="PromQL compatibility in Elasticsearch rising from zero to 80% between 9.4 Tech Preview and 9.5 GA" /><p>This post is about how we built it.</p><p>Key takeaways:</p><ul><li><p><strong>One engine:</strong> The implementation combines Elasticsearch’s mature distributed planning, storage, and testing infrastructure with its newer compute engine, which provides a columnar execution runtime. This lets PromQL reuse proven Elasticsearch capabilities while executing through a modern, native vectorized pipeline rather than introducing a separate runtime.</p></li><li><p><strong>One server:</strong> Elasticsearch implements the Prometheus remote write and query APIs directly, so Prometheus-compatible ingest and queries run without any additional plugins.</p></li><li><p><strong>Engineered for efficiency:</strong> Supporting PromQL required new engine primitives for range-aligned evaluation grids, backward-looking windows, dynamic label grouping, pipeline result reshaping, and compact wide aggregation keys. These primitives allow PromQL queries to execute efficiently end to end, with the relevant semantics implemented directly in the compute engine rather than through external post-processing.</p></li><li><p><strong>Compatibility measured in real use:</strong> In addition to Prometheus compliance tests, we built a differential-testing and quality-control pipeline over 2,000 PromQL queries collected from public repositories. </p></li></ul><p>Read more:</p><ul><li><p><a href="https://www.elastic.co/observability-labs/blog/elasticsearch-supports-promql">Query Prometheus Metrics in Elasticsearch with PromQL</a></p></li><li><p><a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch">Ship Prometheus Metrics to Elasticsearch with Remote Write</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elasticsearch-native-prometheus-api">Bringing Fire to Elasticsearch: Adding Native Prometheus APIs</a></p></li></ul><h2><strong>Why run PromQL on Elasticsearch</strong></h2><p>Many teams already store logs and traces in Elasticsearch while running metrics in Prometheus or another dedicated metrics back end.</p><p>Prometheus and its ecosystem are strong and widely adopted, but large deployments can also bring operational sprawl and scaling challenges, along with limited retention. </p><p>So we set out to combine Elastic’s highly optimized time-series database (TSDB) with a best-in-class metrics ecosystem. The result is a smaller observability stack, with fewer systems to operate and metrics storage that scales horizontally and supports long-term retention.</p><h2><strong>One engine: PromQL and ES|QL share the same compute engine</strong></h2><p>We made an early architectural decision not to run a separate PromQL engine next to Elasticsearch.</p><p>PromQL is instead another front end to the Elasticsearch compute engine.</p><p>This puts PromQL in the normal Elasticsearch development lifecycle. It uses the same planner, distributed execution engine, testing infrastructure, and release process as Elasticsearch itself.</p><p>To learn more about Elasticsearch’s query engine, check <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-metrics-engine-30x-faster-prometheus">our blog</a>.</p><p>Like ES|QL’s time series queries that use the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts"><code>TS</code></a> source command, PromQL is translated into a highly optimized query plan and executed across the cluster. Nodes process columnar batches through vectorized operators, while partial results move through exchanges until the final result is assembled. </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcf3a413e89209805/6aa392b38406d9813bcaa1d8/2.png" alt="PromQL and ES|QL frontends feed one shared Elasticsearch planner and columnar execution DAG across shards" /><p>This also means that PromQL and ES|QL operate over the same execution engine and time-series data. ES|QL can additionally extend a PromQL computation with post-processing that PromQL doesn’t support, such as lookup joins and inline aggregations.</p><p>For example, assume Prometheus request counters are stored in <code>metrics-*</code> and keyed by the <code>service</code> label. A lookup index, <code>service_registry</code>, maps each instance to its owning team and environment and to its service tier:</p><p></p><p></p><p>This architecture requires the execution engine to support PromQL semantics natively and efficiently rather than ES|QL syntax sugar. The following sections describe the changes and new execution primitives we introduced to achieve that.</p><h2><strong>One server: Prometheus remote write and HTTP API built into Elasticsearch</strong></h2><p>Query execution is only half of the story. The Prometheus ecosystem also expects familiar ingest and query APIs.</p><p>Prometheus protocols are the de facto standard for everything metrics in almost every team’s observability stack. So we built the HTTP API directly in the Elasticsearch server, which eliminated the need for a third component and tightened integration stability and performance.</p><p>On the ingest side, we <a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">added</a> an endpoint for the <a href="https://prometheus.io/docs/specs/prw/remote_write_spec/">Prometheus remote write</a> protocol. It accepts Snappy-compressed Protocol Buffer messages, maps labels to TSDS dimensions, maps the metric name/value into metric fields, infers counter versus gauge mappings, and writes directly into TSDS. The built-in template is dynamic, so users don’t have to predeclare every Prometheus label or metric.</p><p>On the query side, Elasticsearch <a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">exposes</a> Prometheus query APIs. A request enters through the Prometheus endpoint and is executed in the compute engine.</p><h2><strong>Running PromQL efficiently in a columnar engine</strong></h2><p>Sharing an execution engine doesn’t mean treating PromQL as syntax sugar over ES|QL. PromQL has different time, grouping, and response semantics, as well as workload characteristics that matter at scale. Supporting it efficiently required extending the compute engine rather than compensating in the API layer.</p><h3><strong>PromQL time grids: Aligning evaluation steps with TSTEP</strong></h3><p>Time-series query engines are optimized for grouping over time.</p><p>Elasticsearch normally groups timestamps with <code>TBUCKET(...)</code>, which truncates each timestamp to a fixed interval boundary. Truncation is cheap and produces deterministic bucket boundaries. It also makes intermediate results easier to reuse.</p><p>Prometheus defines evaluation points differently. For a range query, timestamps are laid out as fixed steps anchored to the query range, rather than derived by truncating each sample timestamp. Two queries with the same step but different range boundaries can therefore produce different evaluation grids.</p><p>To preserve these semantics, we introduced <code>TSTEP(...)</code>, which derives its grouping grid from the query range and step,  rather than truncating timestamps to globally aligned boundaries.</p><p>PromQL uses <code>TSTEP(...)</code> internally, preserving Prometheus timestamp semantics while still lowering the operation to a native Elasticsearch execution primitive.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt437155bca2029671/6aa3937927a5312436dcbf38/3.png" alt="TSTEP vs TBUCKET in Elasticsearch: PromQL step grid anchored to query start, TBUCKET to fixed boundaries" /><p>Some might see this as a simple problem, but the nuances matter. </p><p>Take, for example, a query that finds a 5m rolling average of a metric: </p><p></p><p></p><p>At evaluation time <code>T</code>, the result represents the average over the preceding five-minute range: </p><p><code>(T - 5m, T]</code></p><p>When the query is executed with a five-minute step, each output value is labeled with the upper end of its corresponding five-minute window.</p><p>Elasticsearch previously lacked this semantic and supported only forward-looking window aggregation functions:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt97aac2fa95cf657d/6aa393b71ade64390142ee2e/4.png" alt="Forward-looking window aggregation where each bucket covers the interval from timestamp T to T plus W" /><p>We rewrote the window-evaluation path so that both ES|QL and PromQL use a common backward-looking windowing implementation:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1a12ceb27d9ab95b/6aa393dcf08ee141fe856719/5.png" alt="Backward-looking PromQL window covering T minus W to T, the range used by rate and avg_over_time" /><p>Together, <code>TSTEP(...)</code> and backward-looking windows preserve the two time semantics that matter for PromQL range evaluation.</p><h3><strong>Dynamic label grouping: How PromQL </strong><strong><code>without()</code></strong><strong> resolves at runtime</strong></h3><p>Time grids determine when a PromQL expression is evaluated. Aggregation determines which input series are combined and which labels identify each output series.</p><p>For most analytical query engines, that identity is known when the query is planned. The planner can allocate grouping columns and choose an aggregation strategy. It also carries a fixed key through the execution pipeline.</p><p>That’s how ES|QL works:</p><p></p><p></p><p>The output series are grouped by an explicit key <code>(cluster, namespace)</code>.</p><p>PromQL can express the same operation in the opposite direction:</p><p></p><p></p><p>Now we know which dimensions <em>not</em> to use. We don’t necessarily know the full grouping key until the query is executed. This is a small language difference with significant execution consequences. </p><p>One possible implementation is to discover every label used by the metric, subtract <code>instance</code> and <code>pod</code>, and rewrite the expression into an ordinary <code>by(...)</code> aggregation. That adds a <a href="https://www.elastic.co/search-labs/blog/esql-metrics-info-ts-info-time-series-catalog">discovery phase</a> before planning. It also becomes inefficient for high-dimensional metrics where only a subset of all possible dimensions may have useful value in a particular series. Most queries need only a small subset of the available dimensions, so carrying the entire dimension universe as an aggregation key wastes memory and adds bookkeeping overhead.</p><p>We instead extended the time-series execution path with dynamic grouping columns. The engine loads dimensions as the series are read and applies the exclusions per time series. This avoids making the grouping schema a prerequisite for planning and avoids carrying a large sparse set of grouping columns through aggregation.</p><h3><strong>Dimension packing: Keeping wide PromQL grouping keys cheap</strong></h3><p>The <a href="https://prometheus.io/docs/practices/rules/#aggregation">idiomatic</a> way of writing PromQL aggregations involves heavy use of <code>without(...)</code> over <code>by(...)</code>:</p><p></p><p></p><p>Excluding labels rather than explicitly listing them makes dashboards and alerts resilient to schema evolution. If a new label is added to the metric, the query continues to preserve it unless it’s explicitly excluded.</p><p>The consequence for the engine is that effective grouping keys can be wide. Many of those labels often have low cardinality, yet each still participates in every aggregation stage.</p><p>In a columnar engine, each grouping column is normally represented as a separate vector. Ten grouping labels therefore mean 10 vectors flowing through every aggregation operator:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2f491d0d9fdb82f2/6aa3940027a53105ccdcbf3c/6.png" alt="Elasticsearch columnar page: rows split into typed blocks with delta and ordinal dictionary compression" /><p>Dimension fields are declared in the index mapping; the planner knows the schema up front, and grouping keys stay narrow and predictable.</p><p>In the columnar engine, each grouping label is carried as a separate vector or block. A key with 10 labels therefore requires 10 vectors to be read, hashed, compared, and retained by aggregation operators. As key width grows, so does the amount of data and bookkeeping that must move through the pipeline:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt13b25c686bbb46e0/6aa39429de2395f2e2e86d1d/7.png" alt="PromQL aggregation without packing: each grouping label hashed as a separate block into 64-byte keys" /><p>To avoid paying that per-column cost for every additional label, we introduced dimension packing. Before aggregation begins, the engine encodes the full grouping key into a single compact representation. Hash and comparison operations run on the packed key rather than on each block independently:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5a75855ef51c017b/6aa394468406d95667caa1de/8.png" alt="Dimension packing in Elasticsearch encodes PromQL grouping labels into one 16-byte key before hashing" /><p>Packing lets hashing and comparison operate on a single compact key rather than an increasing number of grouping blocks, making aggregation overhead less sensitive to key width. Because the engine is shared, ES|QL time-series queries will benefit from this optimization as well.</p><h3><strong>Building the Prometheus HTTP API response inside the pipeline</strong></h3><p>Unlike Elasticsearch's ES|QL column-oriented response format, the Prometheus response is row-oriented.  The Prometheus API returns one result row per time series, with its samples represented as timestamp-value pairs. </p><p>To support a compatible API layer, we had to regroup in the HTTP layer converting the columnar results into boxed row objects and accumulate them in map- and list-based structures until the complete Prometheus response could be produced. </p><p>We replaced this with the <code>TimeSeriesCollapse</code> compute operator. It groups rows by series and aligns samples to the query’s fixed step grid. It emits the reshaped result as ordinary columnar pages containing one row per series, with aligned multi-valued timestamp and value blocks. And it preserves compact, vectorized block representation throughout the pipeline: </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt99e2ee3137bc2fa6/6aa39496d29b4e02e81da7b6/9.png" alt="TimeSeriesCollapse operator reshapes five columnar rows into two Prometheus time series per output page" /><p>The HTTP layer can now serialize those blocks directly, avoiding the maps, lists, boxed objects, and associated allocations required by the earlier implementation.</p><h2><strong>Testing PromQL compatibility against 2,000 real queries</strong></h2><p>Prometheus <a href="https://github.com/prometheus/compliance">compliance tests</a> were our starting point.</p><p>Even though they gave us a strong baseline, they didn’t tell us how frequently individual PromQL features appear in real workloads. To complement that baseline, we built a second test corpus from over 2,000 PromQL queries collected from public repositories. </p><p>We then classified those queries by the language features and expression patterns they exercise:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf008821c47e46040/6aa394b68406d90c5bcaa1e2/10.png" alt="PromQL feature use across 2,000 real queries: aggregations 59.67%, selectors 57.55%, rate functions 45.21%" /><p>For each compatible query shape, we run the same query against Elasticsearch and Prometheus and compare the results. </p><p>In addition to that, we actively rely on <a href="https://en.wikipedia.org/wiki/Fuzzing">fuzz testing</a>, which catches issues that unit tests alone are unlikely to expose, including differences in timestamp alignment, label retention, aggregation behavior, range-vector evaluation, and response encoding.</p><h2><strong>Which PromQL functions and APIs are supported in 9.5</strong></h2><p>Since 9.4 (technical preview), PromQL support in Elasticsearch has expanded substantially. In Elasticsearch 9.5, both the<a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/promql"> <code>PROMQL</code></a> command in ES|QL and the<a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api"> Prometheus HTTP APIs</a> are generally available (GA), with more than 80% of the PromQL workflows in our real-world corpus now running without modification:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6706c16409dccfd7/6aa394d327a5312acfdcbf41/1.png" alt="PromQL compatibility in Elasticsearch rising from zero to 80% between 9.4 Tech Preview and 9.5 GA" /><p>The main additions since technical preview are:</p><p><strong>Feature</strong></p><p><strong>Example</strong></p><p><strong>Status</strong></p><p>Prometheus remote write ingest</p><p><code>POST /_prometheus/api/v1/write</code></p><p>GA in 9.5</p><p>Range queries</p><p><code>/api/v1/query_range</code></p><p>GA in 9.5</p><p>Instant queries</p><p><code>/api/v1/query</code></p><p>GA in 9.5</p><p>Metric metadata and build info</p><p><code>/api/v1/metadata</code>, <code>/api/v1/status/buildinfo</code></p><p>GA in 9.5</p><p>Native histogram functions</p><p><code>histogram_quantile</code>, <code>histogram_count</code>, <code>histogram_sum</code></p><p>GA in 9.5</p><p>Per-selector offset modifiers</p><p><code>[5m] offset 1h</code></p><p>GA in 9.5</p><p>Top-level <code>or</code> operator</p><p><code>rate(a[5m])</code> or <code>rate(b[5m])</code></p><p>GA in 9.5, up to eight operands</p><h3><strong>Prometheus remote write ingest</strong></h3><p>Elasticsearch <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/tsds-ingest-prometheus-remote-write">accepts</a> Prometheus remote write (v1) messages directly:</p><p></p><p></p><p>Snappy-compressed Protocol Buffer messages are decoded, and labels are mapped to TSDS dimensions. Metric names and values are written into the time-series index. The built-in template is dynamic, so users don’t have to predeclare every Prometheus label or metric.</p><h3><strong>Range and instant queries through the Prometheus HTTP API</strong></h3><p>Both range and instant query endpoints are <a href="https://www.elastic.co/search-labs/blog/elasticsearch-native-prometheus-api">available</a>:</p><p></p><p></p><p></p><p></p><p>Range queries return matrices evaluated over a time window, and instant queries return vectors evaluated at a single timestamp. These endpoints can be used by Kibana, Grafana, or Prometheus-compatible alerting tools, and custom dashboards.</p><h3><strong>Metric metadata and build info endpoints</strong></h3><p>Elasticsearch <a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api#promql-http-api-metadata">exposes</a> metadata about available metrics and a build-info endpoint:</p><p></p><p></p><p></p><p></p><p>The metadata endpoint returns metric types and help text, and the build-info endpoint returns the Prometheus-compatible server version. Grafana and other tools use these endpoints for feature detection and UI behavior.</p><h3><strong>Native histogram functions: histogram_quantile, count, and sum</strong></h3><p>Elasticsearch supports the <a href="https://www.elastic.co/docs/reference/query-languages/promql/functions/histogram">main PromQL operations</a> over native histograms:</p><p></p><p></p><p></p><p></p><p>Native histograms adapt their bucket layout to the data, providing useful precision across a wide value range without requiring users to configure every bucket boundary in advance. Classic histograms continue to work alongside native histograms.</p><h3><strong>Per-selector offset modifiers in PromQL</strong></h3><p>Offset modifiers shift a selector’s time window backward:</p><p></p><p></p><p>This returns the request rate from one hour earlier. Per-selector offsets are commonly used to compare current traffic, latency, or resource usage with an earlier baseline, such as the same period one week ago.</p><h3><strong>Top-level </strong><strong><code>or</code></strong><strong> operator in PromQL</strong></h3><p>Elasticsearch supports the top-level PromQL <code>or</code> operator:</p><p></p><p></p><p>In PromQL, <code>or</code> isn’t a Boolean operation. It performs a union between two sets of time series. Results from the left side are retained; a series from the right side is added only when its label set doesn’t match a series already returned by the left side. This is useful during migrations where the same logical metric may exist under an old and a new name.</p><p>The implementation follows Prometheus’s left-side precedence rules and preserves the <code>__name__</code> label. Top-level chains of up to eight operands are supported.</p><h2><strong>PromQL features not yet supported in Elasticsearch</strong></h2><p>GA doesn’t mean complete PromQL compatibility. Some less common and more complex parts of PromQL remain unsupported. These gaps now define the next phase of the work: </p><p><strong>Feature</strong></p><p><strong>Example</strong></p><p><strong>Status</strong></p><p>Advanced vector matching</p><p><code>on(instance) group_left</code></p><p>Planned</p><p>Sorting and ranking</p><p><code>topk</code>, <code>bottomk</code>, <code>limitk</code>, <code>sort</code>, <code>sort_desc</code></p><p>Planned</p><p>Label manipulation</p><p><code>label_replace</code>, <code>label_join</code></p><p>Planned</p><p>Absolute time modifier</p><p><code>@ 1710000000</code></p><p>Planned</p><p>Mixed-offset compound expressions</p><p><code>rate(...) - rate(... offset 1h)</code></p><p>Planned</p><p>Alerting and target endpoints</p><p><code>/api/v1/alerts</code>, <code>/api/v1/targets</code></p><p>Out of scope</p><h3><strong>Advanced </strong><a href="https://prometheus.io/docs/prometheus/latest/querying/operators/#group-modifiers"><strong>vector matching</strong></a><strong> with </strong><strong><code>on()</code></strong><strong> and </strong><strong><code>group_left</code></strong></h3><p>Some binary operations that require Prometheus vector matching aren’t yet part of GA.</p><p>For example, this query divides per-instance request rates by a per-instance capacity metric:</p><p></p><p></p><p>The <code>on(instance)</code> clause specifies which labels identify matching series. <code>group_left</code> permits many request-rate series to match a single per-instance capacity series, while retaining the labels from the higher-cardinality left-hand side.</p><p>These expressions are common when joining a detailed metric with metadata or a lower-cardinality capacity metric. Basic binary expressions are supported where applicable, while the remaining vector-matching forms are planned work.</p><h3><strong>Sorting and ranking: </strong><strong><code>topk</code></strong><strong>, </strong><strong><code>bottomk</code></strong><strong>, and </strong><strong><code>sort</code></strong></h3><p>Prometheus sorting and ranking functions are also not yet part of GA:</p><p></p><p></p><p>This returns the 10 services with the highest request rate. Similar queries are widely used in “top offenders” dashboards for traffic, latency, errors, and resource consumption.</p><p>The remaining functions include:</p><p></p><p></p><p></p><p></p><p></p><p></p><h3><strong>Label manipulation with label_replace and label_join</strong></h3><p>PromQL can construct or rewrite labels during query evaluation. These functions are particularly useful when dashboard variables, naming conventions, or label schemas don’t match exactly:</p><p></p><p></p><p>This creates an <code>environment</code> label from the <code>cluster</code> label.</p><p>Another common example combines existing labels into a display-oriented label:</p><p></p><p></p><p>This produces a <code>target</code> label, such as <code>payments/api-7f6d9</code>. <code>label_replace(...)</code> and <code>label_join(...)</code> aren’t yet included in GA.</p><h3><strong>Advanced time modifiers: The </strong><strong><code>@</code></strong><strong> modifier and mixed offsets</strong></h3><p>Several advanced time modifiers and expression forms remain outside the GA scope.</p><p>For example, an absolute <code>@</code> modifier evaluates a selector at a fixed Unix timestamp rather than at the query’s normal evaluation time:</p><p></p><p></p><p>This is useful for comparisons against a fixed historical point.</p><p>PromQL also permits expressions in which the two sides use different offsets:</p><p></p><p></p><p>This compares current traffic with traffic one hour earlier. Per-selector <code>offset</code> is available in GA, but not every combination of offsets and compound expressions is part of GA yet.</p><h3><strong>Prometheus API endpoints not yet implemented</strong></h3><p>In addition, the Prometheus HTTP API surface isn’t yet fully complete. Notably:</p><p>Alerting metadata through:</p><p></p><p></p><p>used by tools that inspect active alert state.</p><p>Target discovery through:</p><p></p><p></p><p>used to inspect scrape targets, health, and labels.</p><p>These endpoints concern Prometheus server and scrape-target state rather than querying metrics stored in Elasticsearch.</p><p>For the full list of limitations, see the <a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-limitations#promql-limitations-form-post">PromQL limitations</a> page. </p><h2><strong>Try PromQL in Elasticsearch 9.5</strong></h2><p>To query Prometheus metrics in Elasticsearch 9.5 or Serverless, see the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/promql">PromQL documentation</a> and the <a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api">Prometheus HTTP API reference</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/promql-elasticsearch-compute-engine</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/promql-elasticsearch-compute-engine</guid>
    <category><![CDATA[Query Languages]]></category>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Sergey Sidorov,Felix Barnsteiner]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf690e82ea51bbeec/6aa37da41ade6445cc42eded/unnamed.png" length="0" type="image/png"/>
    <pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Trust, but benchmark: How we let an AI agent optimize Elasticsearch]]></title>
    <description><![CDATA[We share how we built a harness that automatically identifies and implements optimizations in the Elasticsearch codebase.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch executes a diverse set of workloads, including sustained heavy index building and real-time search and analytics. Delivering excellent performance across the board requires going broad in coverage while simultaneously diving deep enough into the codebase to understand optimization opportunities for each workload. Traditionally, human attention has been the bottleneck in this process; there simply aren't enough engineering hours to scrutinize every hot code path looking for inefficiency across a large and evolving surface area.</p><p>However, with the rapid progression of coding agents, performance optimization has become a task we can tackle semiautomatically. Unlike many software engineering challenges, optimizing code offers a cheap and objective verifier. If you ask an AI model to make code faster, there’s a hard number at the end telling you exactly what happened, backed by profiling tools that explain why. This makes performance a perfect candidate for automation, provided you can actually trust the numbers.</p><p>If you simply point a coding agent at a benchmark, you typically get low signal-to-noise: wins that fall inside the variance of the environment, or variations caused by thermal throttling rather than better code. To capture optimizations that actually benefit Elasticsearch users, we had to bridge the gap between "checkable in principle" and "checked in practice." We built a highly trustworthy measurement loop: a <a href="https://en.wikipedia.org/wiki/Agent_harness">harness</a> that assumes the agent will be wrong a good fraction of the time but reliably catches and proves it when it’s right and then helps guide it where to look next.</p><p>Once the machinery is in place, the results speak for themselves. By letting this harness loose on the codebase, we've already begun uncovering meaningful wins across the stack. In part 2 of this post, we’ll dive into some examples it has found so far, including string conversion inefficiencies in Elasticsearch Query Language (ES|QL), an improvement to our NEON vector dot product implementation, and an upgrade opportunity for the gzip library we were using. In this part, we’ll take a look at the design choices we made and how they relate to the broader topic of effective harness development.</p><h2>The AI code optimization pipeline architecture</h2><p>The first step in any software engineering problem is to identify the correct high-level components. We made an architectural choice that turned out to be very helpful for this problem: separate understanding where opportunities exist from the loop making code changes. The agent starts with a real workload but only uses it to mine information about where to seek performance improvements. At this stage, it’s instructed to go broad and consider a range of performance-related signals. Once it has found and classified the hot spots, the agent reads the context of the code around them to understand the optimization opportunities. We use a separate task to condense the ranked list of hot spots into artifacts that a loop can iterate against in minutes: a microbenchmark that we prove exercises the hot path in its real operating regime. Finally, we use a proposer-verifier loop to actually make changes to the codebase to improve performance on the benchmark. This hands off to validation to assess the impact on real workloads at the end. Our CLI (<code>atune</code>) supplies the tools this process needs, and the rest is largely automated by a set of task-specific instructions.</p><p>For context, our high-level architecture looks like the following. Pink boxes are the humans, and teal boxes are the agent. There are three task types, one skeleton loop, and one referee.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0e561ef2d3b61421/6aa3e83fd1556f0fcc7269f8/unnamed.png" alt="AI code optimization pipeline: exploration to benchmark, human approval, exploitation, validation and PR review" /><p><em>An exploration task profiles a real workload and produces ranked opportunities; a human promotes one into an exploitation task, which iterates against an approved microbenchmark and commits each accepted experiment; a validation run on the real workload guards the result before a human reviews and opens the PR. Where no benchmark covers the hot path, a benchmark task authors one and a human approves it into a registry. A performance atlas informs every task and accumulates what each one learns.</em></p><h2>Why performance optimization suits autonomous agents</h2><p>Three properties make a task ideally suited for autonomous work, and it's worth being explicit about them because they provide a checklist you can use to evaluate automation candidates. You want:</p><ol><li><p>An objective verdict so that the agent can be held to something other than its own opinion.</p></li><li><p>A dense guiding signal so that it knows where to look next instead of guessing.</p></li><li><p>A bounded blast radius so that being wrong is affordable.</p></li></ol><p>Performance gives you all three. Benchmarks provide the verdict and profilers provide the gradient, while a rejected patch costs you wall-clock time rather than correctness. The change gets reverted, and the reason gets recorded. Life goes on. Regarding the first two, we've come to think the gradient matters more than the verdict. ″This got faster″ is binary, whereas a profile hints at what to try next. An agent can generate its next hypothesis conditioned on a rich guiding signal, the richer the better, rather than grinding through a list.</p><h2>Signals are what the agent gets to see</h2><p>A useful mental model is that the CLI is the agent's sensory apparatus. That changes how you design each command. Rather than exposing a capability, you design it to return a clear and concise answer to a question about the task at hand, and, when relevant, an explanation the model can reason over, instead of raw data it has to parse and interpret. These are the signals that the agent acts on, and we ended up with the following for our harness:</p><p><strong>Signal</strong></p><p><strong>Question it answers</strong></p><p>Facet-decomposed macro profile</p><p>Where in the code does real workload time go, per query type?</p><p>Allocation and lock sampling, in the same capture</p><p>Is the cost cycles, garbage, or contention?</p><p>Cost-composition classification</p><p>Is this in scope compute, other product code, GC, JIT tax, or parked threads?</p><p>Input-shape instrumentation</p><p>What does the workload actually feed this code?</p><p>Statistical verdict</p><p>Did this change help, at this measured noise floor?</p><p>Allocation-rate comparison</p><p>Did the new code end up allocating more?</p><p>Interpreted disassembly</p><p>Why did that result happen?</p><p>End-to-end A/B guard, with differential profile attribution</p><p>Did anything appear to break, and was it us?</p><p>Environment check</p><p>Is this machine even fit to measure right now?</p><p>Upstream duplicate search</p><p>Has somebody already reported or fixed this?</p><p>Four of these signals are worth dwelling on, because in each case the tool encodes a judgment that the agent would otherwise have had to keep making by hand.</p><p>Facet-decomposition is the clearest example. Blended CPU shares hide breadth: if you profile a mixed query workload, the grouping hash map insert and the percentiles sketch update can both show up as single-digit percentages of the total run time and look comparable. They aren't comparable at all, because the hash map insert is paid for by nearly every aggregation query, while the sketch update is only paid for when somebody asks for percentiles. So the profiler runs each named query facet as its own race, and every opportunity the agent records carries a breadth field (universal, broad, or narrow) and gets ranked by headroom × tractability × breadth. A universal 3% beats a narrow 10%. Putting the ranking function in the tool prevents it from having to be rediscovered on every run.</p><p>The cost-composition classification works as a router. Rather than handing the agent a flat top-N list of frame names, it buckets every sampled stack into ″in scope compute,″ ″other product compute,″ ″GC,″ ″JIT and safepoint overhead,″ ″off CPU waiting,″ and ″parked threads.″ Each of those buckets implies a different kind of investigation. If GC is above about 15%, the real target is allocation rate, and the CPU top frames will actively mislead you, because they show where objects were collected rather than where they were created. If JIT and safepoint overhead is above about 25%, you're looking at a ceiling rather than an opportunity, since no in-scope code change will move it. If threads are parked and core utilization is low, this is a concurrency problem and CPU flame graphs are the wrong instrument entirely. We wrote that mapping into the playbook as a table, so the model (even a cheap one) reads a profile the way that an experienced engineer would, rather than reaching straight for the top frame. The classification is a prior for forming a hypothesis, though, not a substitute for evidence, so the agent still has to cite specific frames when it proposes an experiment.</p><p>Interpreted disassembly is a tool we hadn't originally provided, but it most definitely earns its place. It helps answer <em>why</em>, the question that unblocks the next hypothesis. Flame graphs tell you where the time goes; they rarely tell you why a change made things worse. So <code>atune asm</code> runs the benchmark briefly with the JIT told to print the assembly for one hot method, captures both sides of the working-tree diff, reduces each to the final C2 compilation, normalizes the addresses, and diffs them. The diff alone would likely still be 4,000 lines of aarch64, so on top of it sits an interpretation layer: a per-mnemonic delta, a net instruction count, the compilation tier that was actually captured, and a vectorization signal that counts vector register references on each side and raises a warning when they're eliminated, halved, or narrowed from <a href="https://en.wikipedia.org/wiki/Advanced_Vector_Extensions">Advanced Vector Extensions</a> (AVX) to <a href="https://en.wikipedia.org/wiki/Streaming_SIMD_Extensions">Streaming SIMD Extensions</a> (SSE) width. The playbook then maps mnemonic patterns to causes:</p><p><strong>Pattern in the diff</strong></p><p><strong>Likely cause</strong></p><p><code>b.eq</code><code>/</code><code>b.ne</code> up, <code>csel</code> down</p><p>New unpredictable branches</p><p>Clusters of <code>str</code><code>/</code><code>ldr</code> against the stack pointer</p><p>The compiler ran out of registers</p><p>NEON loads replaced by scalar compares</p><p>The vector path degraded</p><p>In one experiment, the agent fused two <a href="https://en.wikipedia.org/wiki/Single_instruction,_multiple_data">SIMD</a> mask extractions into one, and the benchmark regressed by 26%. The vectorization warning explained it in about 10 seconds. Without that tool, the agent has a dead end and no working model of the machine; with it, it has a corrected model and several new ideas.</p><p>The fourth signal is less a single tool than a habit; the instruments check themselves. Core utilization is derived two independent ways, from sample density and from process sampling, so the two can be compared. The classification is rejected if the unclassified bucket exceeds a budget, on the grounds that a breakdown which can't account for its own samples shouldn't be reasoned over. The disassembly capture warns when the compilation it caught isn't the steady-state one. The upstream duplicate search is restricted to read-only commands, and that restriction is enforced by a test that greps the source, so no future edit can quietly reintroduce the ability to file anything. Each of these exists because a tool that can be confidently wrong is worse than a tool that is merely absent.</p><p>One small piece of design is worth highlighting as a specific instance of good return practice. <code>atune compare</code> returns 0 for improved, 1 for no change, 2 for regressed, and 3 for error, and the loop branches on this code. That means no parsing and no ambiguity about what the verdict was. Plus, no tokens are spent interpreting prose.</p><p>If you take one thing away from our CLI design, it’s a broader design principle. In a general setting, the interesting thing isn't the individual signals we found useful to understand performance; it's that the CLI is capturing and packaging the judgment of an experienced performance engineer into tools that return answers rather than raw data. Structurally imposing good judgment about the problem an agent is tasked with improves outcomes. The right CLI is as much part of that story as the instructions. Furthermore, tokens are saved by tools that return decisions and digests rather than data. That means a comparison verdict instead of raw JMH output; a triage summary sitting on top of a 4,000-line disassembly diff.</p><h2>From 20-second probes to hours of validation</h2><p>Building a verifier that you can afford to consult is a separate problem from building one that you can trust. This covers the affordability part. Or, if you like aphorisms, real workloads are where truth lives and where iteration goes to die. A macro profile takes 45 to 60 minutes, and an end-to-end validation run takes hours. But a microbenchmark takes minutes. That's why the exploration-then-exploitation split works; you go broad on the real workload once and then hand off to a microbenchmark that you can iterate against in minutes.</p><p>The catch is that the handoff is only sound if the microbenchmark exercises the hot path in a realistic operating regime; that is, with the right cardinality and right data distribution. If you get that wrong, your fast loop spins fast but in the wrong direction. Until we finalized the handoff procedure, we saw cases where the agent accepted changes on a benchmark whose key distributions happened to flatter it, and only the end-to-end run caught the problem.</p><h3>Handing off from exploration to exploitation</h3><p>The handoff between the two phases is structured rather than informal. An exploration task's primary deliverable is a set of opportunity records, and each one carries the scope paths that an exploitation task would be allowed to edit, a headroom estimate, a classification (constant factor, structural,</p><p>allocation, or concurrency), and the benchmark it would gate on (or an explicit "no benchmark coverage" flag, if none exists). Each also carries a narrow test pattern so that the correctness gate stays cheap. The test pattern field exists because of a specific incident; an exploration task omitted it, and the resulting exploitation task ran a very heavy test suite on every experiment. The fix was to change the upstream artifact rather than add an instruction downstream. This is a pattern we use repeatedly; make and record decisions as early as possible rather than re-derive them each time.</p><h3>Validating a new benchmark before it can gate anything</h3><p>Where benchmark coverage is genuinely missing, a dedicated benchmark task authors one, and that new benchmark has to pass validity checks before anything can rely on it. Two of them are mechanical: what fraction of the benchmark’s hot self-time comes from frames that actually appear in the production profile and whether the parameters fall inside the input shapes we measured. The third is a checklist that the agent has to attest item by item. It exists to avoid the JIT getting a simpler world than production.</p><ol><li><p>Inputs have to be reshuffled rather than fixed or sorted, so the branch predictor doesn’t get too good.</p></li><li><p>Results have to be consumed, or dead-code elimination deletes the thing that you meant to measure.</p></li><li><p>Inputs must not be compile-time constants, or they get folded away.</p></li><li><p>Call sites have to see roughly the product mix of types, because a monomorphic call site inlines, whereas a megamorphic one doesn’t.</p></li></ol><p>A human then approves it into a hash-pinned registry. Until that happens, it's inert, because task setup refuses any task citing an unapproved benchmark. We think of that approval as the strongest gate in the system, and it's deliberately placed. An approved benchmark can decide accept or reject in every future task, so it's the one place where we ask for a human signature on an artifact rather than on a decision.</p><h3>The validation ladder</h3><p>Underneath all of this sits a hierarchy of feedback mechanisms with the property that each rung is cheaper and weaker than the one below it, and the cheap tiers are for rejection only.</p><p><strong>Tier</strong></p><p><strong>Cost</strong></p><p><strong>Role</strong></p><p>probe</p><p>~20–60 s</p><p>Directionally right? Can never accept</p><p>codegen capture</p><p>~4 min</p><p>Why did that happen?</p><p>screen</p><p>~5–15 min</p><p>Cheap statistical filter</p><p>confirm</p><p>~20–60 min</p><p>The accept decision</p><p>end-to-end</p><p>hours</p><p>Regression guard, advisory</p><p>The asymmetry between accept and reject is doing real work here. A probe is a single paired fork, whereas the accept predicate requires a full confirm run with a matching calibration record. Agents are very good at telling believable stories; indeed, they're trained on many tasks judged by both LLMs and by humans, so being convincing is actively rewarded. We don't want an agent to be able to promote a cheap signal into a decision by being persuasive about it.</p><p>For this task benchmark, wall-clock is a real cost consideration. A confirm run might take an hour. So one has to weigh carefully all the costs involved when choosing the setup. A model that lands one hypothesis in four typically beats a cheaper one landing one in 10 by a margin on end-to-end metrics. The usual instinct to down-spec the model on a long-running loop is exactly backward here. When your loop has an uncertain outcome and significant costs beyond the tokens it consumes, you may well find yourself in the same situation.</p><h2>How do you know a performance improvement is real?</h2><p><em>Is it faster?</em> is a statistical question. So we made the accept predicate code rather than judgment and put it somewhere the agent can't bypass. There are four ideas in the accept decision, and in each case, the alternative we rejected is as informative as the choice we made.</p><h3>Forks are the statistical unit</h3><p>Each <a href="https://github.com/openjdk/jmh">JMH</a> fork collapses to its mean, and verdicts come from an exact <a href="https://en.wikipedia.org/wiki/Mann%E2%80%93Whitney_U_test">two-sided Mann-Whitney U test</a> at α = 0.05 plus a seeded bootstrap confidence interval over three to five fork means per side. The reason is that iterations within a fork share JIT and heap state and are therefore autocorrelated, so treating them as independent samples manufactures significance out of nothing. We rejected comparing single-run scores, which is pure noise, and iteration-level <a href="https://en.wikipedia.org/wiki/Student%27s_t-test">t-tests</a>, which can be confidently wrong. Seeding the bootstrap means that a rerun reproduces the verdict exactly because the agent needs to be able to tell the difference between a result that changed and a result that was never stable.</p><h3>Pair candidate and baseline in time</h3><p>The screen (three forks, optionally over a subset of parameters) exists only to kill bad hypotheses in maybe 10 minutes instead of an hour. The accept decision itself comes from a confirm run that measures candidate and baseline back to back using a stash-flip. In an unstable environment, thermal and background drift only cancels if both sides ran under the same conditions, so an hours-old baseline is really a different experiment.</p><h3>The noise floor is measured, not assumed</h3><p>The minimum effect size that a task will accept has to clear the A/A-calibrated coefficient of variation for that specific benchmark on that specific machine, and both the confirm run and the comparison refuse to proceed without a matching calibration record. A 1–2% improvement on a laptop is indistinguishable from noise, and because the floor varies by benchmark and by JDK, any global constant you pick will be too loose somewhere and too tight somewhere else.</p><h3>The accept rule is composite and deliberately conservative</h3><p>A parameter combination counts as improved only if p &lt; α, the effect clears the calibrated floor, and the confidence interval excludes zero. Overall acceptance then requires that nothing regressed (not the primary benchmarks and not the guards) and that at least one primary combination improved. The headline figure is the <a href="https://en.wikipedia.org/wiki/Geometric_mean">geometric mean</a> of the per-combination speedup ratios, which is always positive and composes across experiments, so a task's cumulative improvement is a meaningful number rather than a sum of incomparable percentages.</p><h3>What must not get slower</h3><p>Guards deserve a note of their own, because they answer a different question to the primary benchmarks; not <em>Did this get faster?</em> but <em>What must not get slower while it does?</em> The task definition lists them separately for that reason, and they're typically the operations and the input regimes that the change isn't aimed at. If you're optimizing insert throughput on a hash table, iteration is a guard and so is a collision-heavy key distribution. A change that improves the common case by weakening the hash function will look good on uniformly distributed keys while catastrophically degrading more adversarial inputs. We know that because it happened; the collision distribution was missing from the matrix that accepted one of our early experiments, and the task now carries a comment telling future readers never to drop it again for a hash-quality-sensitive scope. While guards cost wall-clock on every confirm, they also surface edge-case regressions, and omitting them can be much more costly in the long run.</p><p>Runtime isn't the only thing worth guarding, either. The confirm run also captures normalized allocation rate on both sides and flags any change that buys speed with more than about 15% extra garbage. That one is advisory rather than blocking, because sometimes the trade is the right one. However, it's the kind of regression a purely time-based accept rule would happily wave through but might raise a red flag to an experienced performance engineer with better understanding of the calling context.</p><h3>The end-to-end gate is one-sided and default open</h3><p>The end-to-end gate is a different statistical problem: small n, high noise, and a very strong prior that we should accept based on our microbenchmark results. Our first design treated accept and reject on an equal footing, and it produced multiple clearly spurious rejections, so the redesign is one-sided and default open. An operation is flagged only when the median regression exceeds the threshold and every candidate repetition is slower than every base repetition. That full-separation criterion is the nonparametric one-sided test at this sample size, and it's robust to the single outlier repetition that would occasionally fool us otherwise. A flag then also has to be corroborated against a differential CPU flame graph, where only a rise in the task's own in-scope CPU share counts as real. Near misses get reported for transparency but don't trigger triage, and nothing is ever auto-rejected; a flag is a request for human attention rather than a verdict. We have a final backstop which is the large suite of performance tests that we already run against Elasticsearch on a daily basis.</p><p>The one-sided gate lesson generalizes beyond benchmarking. A noisy gate should be one-sided and default open where there is strong prior reason to accept. A symmetric threshold on a noisy signal doesn't just cost you real wins, it teaches the loop to distrust its own instruments, and that’s a much more expensive failure.</p><h2>Exploration and exploitation need different permissions</h2><p>Exploration and exploitation might look like two phases of one activity, but they have different inputs (a macro workload versus a pinned scope) and different outputs (ranked opportunities versus commits). They also have different failure modes, which means they want different permissions. We made the split a first-class property of a task, which lets us enforce it; an exploration task literally cannot commit. The baselining, comparison, checkpointing, and validation CLI all refuse exploration tasks, benchmarking allows probes only, and every probe diff is always reverted. A broad, speculative survey is safe because nothing it does can edit the code.</p><p>A nice ancillary benefit is prompt focus. Each type reads one playbook, in full, with the others explicitly not loaded. If you try to write a single document covering both "find where the headroom is" and "land a validated win inside this scope", you get something that does neither well, because the instructions for good exploration (follow the profile, widen the net, a broad survey is preferred) are close to the opposite of the instructions for good exploitation (one hypothesis at a time, minimal diff, never widen scope).</p><h2>AI agent memory: Journals, knowledge bases, and postmortems</h2><p>Sessions are ephemeral, but what you can learn from them isn't, so the harness accumulates three durable assets, plus one disposable view derived from them. These are a journal of the code changes we’ve tried, a knowledge base of how the code performs, postmortems of when the harness failed and, because sessions can be stopped and resumed, a session summary. What makes them work together is a clean ownership rule about which kind of fact goes where.</p><p>The journal records what we tried and measured. It's append-only, one file per task, and one record per experiment, and it's written before the code is edited. Rejections carry a forward-looking note in the form "do not retry X because Y", which is probably the highest value line, because it's what stops the next session re-deriving a dead end. Records also carry the environment and the driving model, which means that hypothesis hit rates are comparable across models.</p><p>The knowledge base records how the code works and how it performs. It's an indexed collection of per-area summaries, each stamped with the commit it was written against, and every playbook ends with an upkeep step that appends whatever durable facts the run turned up. Because the performance characteristics of the JDK also change from time to time, for example, a new <a href="https://download.java.net/java/early_access/loom/docs/api/jdk.incubator.vector/jdk/incubator/vector/Vector.html">Vector API</a> might implement vector masking more efficiently on AArch64, findings from profile data are also tagged with the JVM version they apply to.</p><p>The postmortems record mistakes that the agent has made in the past, and they're indexed by symptom rather than by date. The question a session actually has when a number looks wrong is <em>Have I seen this shape of wrongness before?</em>, and a chronological list doesn't answer it. So the rows read like "validation fails on operations structurally unrelated to your diff", or "screen reports no matching combinations".</p><p>Keeping the journal and the knowledge base distinct sounds pedantic but isn't because without the rule, both of them turn into a diary that has a tendency to bloat the context window or miss critical information in context.</p><p>The disposable state is a session handoff, and our advice is to never write it by hand and to avoid asking a model for a session summary, if possible. Our harness regenerates a one-page digest mechanically from the journal, so a resumed session doesn’t spend time and tokens reconstructing where the task had got to. Because it’s derived rather than authored, it can’t drift from the record in the way hand-maintained content does. That’s also why it isn’t part of the audit trail; because it’s cheap to regenerate it from the journal.</p><p>Mechanical session handoffs point to two lessons, and they turn out to be the same one. Every time a person appears in the loop, it’s a source of friction and an opportunity for error. And every time you reach for a model, ask whether code can do the same job. This seems like an odd thing to advocate in a project whose primary premise is delegating to a model, but the habit is easy to fall into once you have one to hand. Judgment is expensive, wherever it happens to sit, so the person and model both have to earn their place on merit.</p><p>Two further things are critical for durable agent memory. The first is that knowledge rots, so you have to lint it. Elasticsearch's main branch moves daily, which means the knowledge base's file citations go stale, so one linter checks them. Another checks that every command, flag, and path cited in the agent-facing docs actually exists, that every postmortem is linked from the index, and that every task type has a playbook, and it runs as part of the test suite. Treating prose written for an agent as a testable artifact is the reason it stays true. The second is that loading discipline is half of memory. The instruction is to load the index and then the one or two summaries matching the task's scope; never to bulk-load the rest. Memory you can't afford to read isn't memory.</p><h2>Coding agent guardrails: Containment, scope, and stop conditions</h2><p>Elasticsearch is millions of lines of code, and an unscoped "make it faster" run against a codebase that size is unreviewable and unfalsifiable. It’s also expensive. The converse is that a verifier only protects what it can see, so we have to apply the same boundary to what it can change. The outcome is a task that fixes its goal, allowed paths, benchmarks, thresholds, and stop conditions before any code is edited, and none of those are things the agent may change during a run.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2bae81c11c37e29b/6aa3e867d1556f62a67269fc/unnamed.png" alt="Coding agent guardrails: human owns scope and thresholds, agent owns judgement, atune CLI enforces mechanically" /><p><em>The human owns the safety envelope, defining scope, thresholds, and stop conditions before any code is edited, and owning every handoff that crosses a trust boundary. The agent session owns judgment: reading profiles, forming one hypothesis at a time, editing inside scope. The atune CLI owns mechanical enforcement: scope checks, correctness gates, statistics, calibration, stop conditions, environment checks, and the signals. Persisted state is the record: a worktree pinned at a base commit, an append-only journal, and generated reports.</em></p><p>Containment happens in two layers. The first layer is coarse and task-agnostic; it’s a single static permissions file lets the session edit the per-task worktree (which includes a local branch of Elasticsearch, all journal entries, and CLI artifacts) and the knowledge base, and it denies the pristine Elasticsearch clone, the task definitions, the harness config, and the audit trail. The second is precise and per-task: a git-level scope check that runs over tracked and untracked files before any build, so an out-of-scope edit is rejected before it can be benchmarked or committed.</p><p>Two containment layers have a nice corollary. Widening the first layer to the whole worktree costs nothing, because nothing out of scope survives the second one. Coarse containment plus precise scope beats trying to make a single mechanism do both jobs, which is what we tried first and which left us with a permissions file that needed editing for every new task.</p><p>Stop conditions are mechanical. There's a maximum number of experiments and of consecutive rejections, along with a cumulative improvement target, and proposing a new experiment is refused once one of them fires. A human can override; the agent can't. That asymmetry is what makes the gates independent of which model happens to be driving.</p><p>Tests are add-only. New test files ship with the checkpoint, and modifying or deleting an existing test is blocked by the scope check. This closes the single most tempting shortcut in the entire problem space by construction rather than by instruction, which seems like the right way to handle any shortcut you'd otherwise have to keep asking an agent not to take.</p><p>The harness lives outside the code it optimizes. The Elasticsearch clone is a separate, gitignored directory, and each task gets a worktree pinned at a base commit. The payoffs compound; the subject repo stays pristine and upstream mergeable with nothing related to the harness leaking into a PR, and the audit trail is versioned independently of a codebase that moves daily. Also, tasks are isolated from each other and from any developer checkout. Targeting a newer Elasticsearch means creating a new task rather than repointing an old one, because that task's numbers are tied to its base. It also means that the harness is retargetable in principle, since the Elasticsearch-specific parts are configuration, knowledge base, and benchmark registry rather than architecture.</p><h3>The decisions the agent never makes</h3><p>The rule we settled on is that the agent runs the loops and a human owns every step that crosses a trust boundary. That is creating work, blessing a measurement instrument, publishing a branch, or acting outside the repo. None of those are in the agent's allowlist, and each has a reason worth stating:</p><ul><li><p>Humans have to sign off on the task because the thing being constrained can't set its own limits.</p></li><li><p>Calibrating the noise floor needs to be done once per benchmark, per machine, and by default is measured rather than assumed. However, it's rather expensive and we allow a human to override if they know the environment well.</p></li><li><p>Deciding what to optimize next by promoting an opportunity is a judgment and a new scope.</p></li><li><p>Since benchmarks go on to gate other tasks, we consider reviewing this artifact part of the correctness safety net.</p></li><li><p>We leave outward-facing actions, such as pushing a branch or filing an issue, to a human until we're confident in the process.</p></li><li><p>We allow actions to be forced, but the override has to sit outside the thing being overridden.</p></li></ul><p>How the human actions get surfaced in the workflow matters. The generated report and the session handoff both print the human actions currently due, at the moment they become due, rather than leaving them to be inferred from the playbook.</p><p>We're deliberately not taking a position on how permanent the manual processes are. The right amount of supervision for a new technology is an empirical question, and we'd rather measure it than argue about it. We’ve started with a relatively high degree of supervision because that's the cheap direction in which to be wrong (a gate you never needed is easier to remove than a regression you shipped) and because the harness makes the question answerable. Every gate is a named, logged transition, so over time we can see which of them ever changed an outcome and which only ever cost friction. In summary, measure first, and then refine.</p><h2>Building the harness is the same kind of loop</h2><p>A lot of the harness design didn't fall out of an initial design document. The signal set, the ranking function, the shape of a task, and the exact wording of a playbook rule each came from watching a run go wrong. If there's one piece of advice here that generalizes, it's to use the thing before it's ready and to instrument your own disappointment.</p><p>The clearest example is a rule we now call <em>distrust surprising results</em>. A validation run reported that every operation had regressed, the worst of them by 14.8%. It was wrong twice over. A target operation pattern had overmatched a completely different code path, and a stale output directory from an earlier run was being read alongside the new one. Offered a coherent story, the agent took it and reverted a change that was actually good.</p><p>What went into the playbook after that incident is not "be careful." It's a three-step check to run before acting on a surprising verdict:</p><ol><li><p>Trace the code path, and confirm that the thing which moved can even reach your diff.</p></li><li><p>Read the raw per-repetition data rather than the summary, and recompute one headline number by hand.</p></li><li><p>Compare the report's shape against a known good run, because a structurally different report implicates the pipeline rather than the code.</p></li></ol><p>Alongside that, there’s another important rule, which is if the agent concludes that the harness is buggy, it must <em>not</em> fix it mid-run, because a mid-run harness change makes every result in that run incomparable. It should journal the evidence and stop.</p><p>Improving the harness is itself a loop worth describing. Asking the model to review its own transcripts and the harness documents, and to propose the rule itself, usually works well. It's good at spotting where its own instructions were ambiguous, in a way that's hard to reproduce by rereading the instructions yourself. What makes that output useful is having somewhere for it to land: a terse rule in the playbook, the narrative in a dated postmortem, a symptom keyed index row, and a linter that keeps the citations honest.</p><p>Restraint turns out to be part of the same discipline. The design document carries an explicit list of extension points that we've deliberately not built, because the need for them is still speculative. That's the same "don't guess, wait for evidence" rule we impose on the optimization loop, applied to ourselves.</p><h2>How this applies beyond performance optimization</h2><p>A few of these themes aren't specific to performance work or to Elasticsearch.</p><p>Verifiable work is the current frontier. The same insight drives <a href="https://arxiv.org/pdf/2411.15124">reinforcement learning with verifiable rewards</a>, and it’s what <a href="https://arxiv.org/pdf/2506.13131">AlphaEvolve</a> is built around. The tasks agents are consistently good at are the ones that come with a cheap oracle (tests, compilers, benchmarks), and the interesting move isn't finding more such domains but manufacturing oracles for domains that lack them. Performance is an instructive case precisely because the oracle is, in some senses, obvious, and yet building it still took most of the engineering.</p><p>The referee pattern generalizes, too. Separating a fallible optimizer from mechanical enforcement is the same shape as sandboxed execution and policy engines: judgment in the model, invariants in code. The practical consequence is that the system's safety doesn't depend on which model drives it. A weaker model wastes benchmark time, but it can't corrupt the code or accept a bogus win.</p><p><a href="https://en.wikipedia.org/wiki/Goodhart%27s_law">Goodhart</a> is a standing adversary for any optimization task that uses an agent, and it has a <a href="https://arxiv.org/pdf/2209.13085">formal treatment worth reading</a>. An agent optimizes <em>exactly</em> what you tell it to, so the two-tier benchmark structure, add-only tests, benchmark approval registry, adversarial guard distributions, and a marker file that makes timing an instrumented build mechanically impossible are all one design theme wearing different clothes.</p><p>Tools are context engineering. The <a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents">consensus is drifting away from "expose everything"</a> and toward a few well-shaped <a href="https://www.anthropic.com/engineering/writing-tools-for-agents">tools that return digests</a>: the principles of progressive disclosure, self-documenting interfaces, and verdicts rather than payloads.</p><p>Memory is becoming architecture. A rules file, per-task playbooks, a durable knowledge base, and an append-only journal form a hierarchy with different lifetimes, owners, and loading rules, and the hard part is eviction and staleness rather than storage; hence, the linters.</p><p>Finally, human-in-the-loop is a dial rather than a switch, so where it should sit is something to measure per domain rather than assert.</p><h2>What's in part 2 of this post</h2><p>The harness design is a set of hypotheses about what autonomous performance work needs, and the harness was built so that we could test them. Part 2 is that test: the first four PRs we raised using it, their gains, how many hypotheses it took to get each one, which gates actually caught something, and where the harness got in its own way. That includes the changes which didn't survive end-to-end validation, since as usual, the rejections are as informative.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/ai-code-optimization-elasticsearch-agent-harness</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/ai-code-optimization-elasticsearch-agent-harness</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Thomas Veasey,Chris Hegarty]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf9efb1bf91e4491b/6aa3e56cd909f868f36d69a6/unnamed.png" length="0" type="image/png"/>
    <pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[No more allocation delays: Decoupling snapshots from shard relocation in stateless Elasticsearch]]></title>
    <description><![CDATA[Clusters scale out under load without waiting for a snapshot to finish, because snapshots now read straight from the object store and no longer pin shards in place.]]></description>
    <content:encoded><![CDATA[<p>Stateless snapshots now read directly from the object store. Server-side "undesired allocation due to snapshot" warnings stopped entirely after the release. Primary shards stay free to relocate while a snapshot runs, so clusters scale out under load without waiting for one to finish. Across the fleet, cache misses dropped by more than 60% and median cache population throughput rose roughly 50%. </p><h2>Why snapshots pin primary shards in stateful Elasticsearch</h2><p>In traditional stateful Elasticsearch, snapshots lock primary shards to their active nodes, completely preventing relocation. That works fine when cluster topology is stable and nodes stay online between maintenance windows, but stateless Elasticsearch works differently. Index data lives in an external object store, with local disk as a cache, and the cluster scales automatically based on CPU, memory, and data size, both vertically (upsizing nodes) and horizontally (adding nodes).</p><p>During vertical scale-up, existing nodes must vacate all shards and shut down before new hardware takes over. Since Elasticsearch version 8.13, shard snapshots can pause during node shutdowns and resume after relocation, so a long-running snapshot doesn't block an infrastructure update.</p><p>Horizontal scale-out is a different story: No nodes shut down, so pause logic never triggers. New nodes sit idle while existing nodes finish their snapshots, and as soon as a snapshot is queued, the primary shard is pinned to its node, significantly delaying relocation.</p><p>Clusters typically scale out because they're already under heavy load. Blocking shard relocations at that moment limits the cluster's ability to reduce pressure, which shows up as degraded indexing throughput and higher latency. The resource imbalance can also trigger unexpected autoscaling behavior. And even when overall topology stays the same, shard locking disrupts hotspot mitigation and workload distribution. These failures used to surface as server-side warnings: "undesired allocation due to snapshot."</p><p>
</p><p><strong>Stateful Elasticsearch</strong></p><p><strong>Stateless Elasticsearch</strong></p><p>Snapshot reads from</p><p>Local shard data on the node holding the primary</p><p>The object store, using file locations recorded in the commit</p><p>Primary shard during snapshot</p><p>Pinned to its node until the snapshot completes</p><p>Free to relocate at any time</p><p>Effect on horizontal scale-out</p><p>New nodes wait for in-flight snapshots before taking shards</p><p>New nodes take shards immediately, regardless of snapshot state</p><p></p><h2>How stateless snapshots read directly from the object store</h2><p>A shard snapshot pins primary shards because it needs to read local shard data. In stateless Elasticsearch, that data already lives in the object store, so reading from local disk is unnecessary. Letting snapshots read directly from the object store removes the requirement to lock primary shards. They can relocate freely, and backup is decoupled from cluster balancing.</p><p>Stateless commits include location information for each data file in the object store, so snapshots can read and stream directly to the snapshot repository (a separate object store bucket). In the future, we plan to look at server-side ranged copies, which object stores support natively, to skip the local copy step entirely.</p><h2>Tracking commits when shards relocate mid-snapshot</h2><p>A snapshot is bound to a specific commit point that determines which files to back up, and those files must remain accessible for the full duration of the operation. In stateful clusters, this is simple: The snapshotting node and the data node are the same, so the commit is managed locally and held until completion.</p><p>In a stateless model, the snapshotting node and the data node can be entirely separate, or they can diverge if a shard relocates mid-snapshot. To handle this, we added a transport action that acquires commits on remote data nodes over the network. The data node tracks which commit belongs to which snapshot and releases it once cluster state signals completion.</p><p>There's a wrinkle during relocation. A stationary shard relies on its commit point to preserve files. A relocating shard must release its commit so its local store can close cleanly. To keep files accessible through that transition, a newly recovered primary temporarily preserves all existing data files in the object store until notified of snapshot completion via cluster state. This handles both graceful relocations and ungraceful recovery from node or engine failures.</p><h2>No more allocation delays and improved cache stats</h2><p>After stateless snapshots shipped, the server-side "undesired allocation due to snapshot" warnings stopped. The chart below shows the before and after, with the release marked by the red arrow. Hotspot mitigation became more responsive because shard relocations no longer had to wait for backup operations.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0c08d57928ea1c6c/6a97a33f2707c544c5329695/unnamed.png" alt="Bar chart showing undesired allocation due to snapshot warnings dropping to zero after stateless snapshots shipped" /><p>Cache use is also improved. Snapshots that bypass local shard data stop competing with indexing for cache space. After the release (also marked in the chart), we observed the following two positive changes in cache metrics:</p><ol><li><p>The median cache population throughput, defined as bytes per second for filling the local disk cache from the object store, increased about 50%.</p></li><li><p>Cache misses, where data must be retrieved from the object store to fill local disk cache, have dropped more than 60%. </p></li></ol><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0b85caa56b35485d/6a97a3683eabd0c326440fab/unnamed_(1).png" alt="Charts showing cache population throughput rising 50% at p50 and cache misses falling over 60% after release" /><h2>What comes after stateless snapshots</h2><p>Object-store-native architectures are increasingly the standard for cloud-native data systems, and stateless snapshots are a step toward fully exploiting that model across Elasticsearch operations. Backups read from the object store, and shards move freely. Neither process waits on the other. Removing the local shard dependency is a step toward further modularizing the stateless architecture.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/stateless-snapshots-shard-relocation</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/stateless-snapshots-shard-relocation</guid>
    <category><![CDATA[Elastic Cloud Serverless]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Operations]]></category>
    <dc:creator><![CDATA[David Turner,Yang Wang]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt685f7746435d0451/6a97a2aaf08ee14b39853cbb/unnamed.png" length="0" type="image/png"/>
    <pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Avoiding and Correcting Hotspots: How Elasticsearch Serverless Balances Shards]]></title>
    <description><![CDATA[Elasticsearch Serverless replaces the Elasticsearch node-weight based shard rebalancing algorithm with resource usage aware rebalancing that avoids index shard colocation, OOM events and write load hotspotting]]></description>
    <content:encoded><![CDATA[<p>The Elasticsearch Serverless Balancer addresses write load hotspots, prevents data node out-of-memory (OOM) events and avoids index-level hotspots in Elasticsearch Serverless clusters: these are workload edge cases that in non-Serverless require manual intervention and custom tuning of cluster settings. Serverless shard balancing focuses on staying within the bounds of node-level resource constraints. Rebalancing moves are explainable, where moves are made explicitly to either avoid performance degradation or correct hotspots when they develop. Shard movements are generally found to be fewer, as well.</p><h2>How Elasticsearch Shard Balancing Works</h2><p>Elasticsearch uses a weights-based algorithm to create a Desired Balance, an assignment of shards to data nodes. The Balancer determines the target allocation of shards across a cluster of nodes using four key metrics weighted in a linear algorithm. A total weight is calculated per node, and the shard balancer aims to equalize the total weights across cluster nodes. A final Desired Balance shard allocation is precomputed based on the latest cluster state information, and then the elected master node initiates incremental shard moves to reach the desired shard allocation.</p><p>The four metrics are:</p><ul><li><p><strong>Write Load:</strong> the total write threadpool activity per node, using the sum of threadpool indexing activity per data-stream shard.</p></li><li><p><strong>Disk Usage:</strong> the total disk usage of shards per node, using the sum of disk space used per shard.</p></li><li><p><strong>Shard Count:</strong> the total number of shards assigned to a node.</p></li><li><p><strong>Index Balance (shard anti-affinity):</strong> per index, how many shards in the index are assigned to the node.</p></li></ul><p>The total weight of a node is calculated using a linear algorithm that finds the deviation from the node-level cluster average for each individual metric, applies a different weight factor multiplier to each, and then takes the sum of all resultant values. The weight factor multipliers attempt to equalize the relative magnitude of each metric so that metrics with large values do not eclipse metrics with naturally small values. Write load tends to be a small value, related to thread usage, and thus gets multiplied by a relatively larger weight factor of <code>10</code>; whereas disk usage in bytes is a very large number and therefore gets multiplied by a tiny weight factor of <code>2e-11</code>.</p><p>The following are the cluster settings with default values, representing the different weight factors:</p><p><code>cluster.routing.allocation.balance.shard: 0.45</code></p><p><code>cluster.routing.allocation.balance.index: 0.55</code></p><p><code>cluster.routing.allocation.balance.disk_usage: 2e-11</code></p><p><code>cluster.routing.allocation.balance.write_load: 10.0</code></p><p>The linear algorithm looks something like this:</p>final float shardWeightFactor =
    settingValue("cluster.routing.allocation.balance.shard");
final float writeLoadWeightFactor = 
    settingValue("cluster.routing.allocation.balance.write_load");
final float diskUsageWeightFactor = 
    settingValue("cluster.routing.allocation.balance.disk_usage");
final float indexWeightFactor = 
    settingValue("cluster.routing.allocation.balance.index");

final float shardCountDeviation = numShardsOnNode - averageShardsPerNode;
final float writeLoadDeviation = totalWriteLoadOnNode - averageWriteLoadPerNode;
final float diskUsageDeviation = totalShardDiskUsageOnNode - averageShardDiskUsagePerNode;
final float indexDeviation = numIndexShardsOnNode - averageNumIndexShardsPerNode;

return shardCountDeviation * shardWeightFactor
    + writeLoadDeviation * writeLoadWeightFactor
    + diskUsageDeviation * diskUsageWeightFactor
    + indexDeviation * indexWeightFactor;<p>Shard movements are triggered to ensure that the difference in total node weight across cluster nodes remains below the <code>cluster.routing.allocation.balance.threshold</code> with a default value of <code>1</code>: whenever the threshold is exceeded, shards are moved from the most heavily weighted nodes to the least heavily weighted nodes until the difference between the most heavily weighted and least heavily weighted node is at or below the <code>threshold</code>. Whenever cluster activity occurs that changes shard allocation (e.g., create/delete index, add/remove node, or the disk usage grows), the Balancer rechecks the weights across nodes and triggers shard rebalancing if the delta between the most and least heavily weighted nodes exceeds the configured threshold. The threshold-based approach attempts to balance the trade-off between keeping the cluster perfectly balanced and minimizing shard movements. Large Elasticsearch deployments that use nodes with greater resources typically benefit from raising the <code>threshold</code> setting: a larger weight delta between nodes reduces shard rebalancing.</p><p>Shard movement is also constrained by strict shard assignment rules that prohibit certain node assignments according to cluster and index level settings. Examples include: not assigning copies of the same shard to the same node or host; not allowing further assignment of shards to a node that does not have spare disk space; and excluding node(s) as host for a particular index. More on this below.</p><h2>How a Balanced Cluster Looks (Based on Weights)</h2><p>Using the linear algorithm and cluster setting defaults previously described, the following is an example of what the Balancer considers balanced. Notably, it can sometimes allow considerable deviation across nodes in any one particular metric. For simplicity, index balance is not included.</p><p><em>Weight Node1 = 0</em>.45 (5 - 6) + 10 (0.3 - 0.33) + 2e-11 (1e+11 - 8e+10) =  <em> - 0.35</em></p><p><em>Weight Node2 </em>= 0.45 (7 - 6) + 10 (0.4 - 0.33) + 2e-11 (4e+10 - 8e+10) = <em>  0.35</em></p><p><em>Weight Node3 </em>= 0.45 (6 - 6) + 10 (0.3 - 0.33) + 2e-11 (1e+11 - 8e+10) = <em>  0.10</em></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt716ecf26a404be16/6a969bee0897906115efbc2d/2.png" alt="Weights-based shard balancing: three Elasticsearch data nodes with 5, 7 and 6 shards, differing disk usage and write load" /><h2>Limitations of Weights-Based Shard Allocation</h2><p>Elasticsearch Serverless deployments are managed by Elastic and could run into many edge cases that, without additional configuration, the weights-based shard allocation handles poorly. In self-managed Elasticsearch deployments, it is possible to work around many of these issues by configuring the cluster to suit the user’s workload. However, Elasticsearch Serverless is configured once and must work across all customer use cases. Issues experienced by some Elasticsearch customers (early adopters of Serverless among them) include:</p><ul><li><p>Continuous rebalancing background noise in active clusters. This could be because the <code>threshold</code> setting needs tuning or because the cluster is very busy.</p></li><li><p>The Balancer’s behavior cannot be tuned in a predictable manner. Adjusting the Balancer settings (individual weight factors) can lead to unpredictable outcomes due to the linear algorithm. For example, decreasing the shard count weight factor relative to the other weight factors can lead to data node OOM events when shard count balancing is deprioritized and too many shards pile up on a single node.</p></li><li><p>No explanation of why the Balancer is making shard moves. The linear algorithm is difficult to understand without relevant node metrics.</p></li><li><p>The linear algorithm allows a high value in one metric to cancel out a low value in another metric. For example, a node can have a higher than average (across cluster nodes) write load, but counterbalance with a lower than average shard count (or vice versa), and the linear algorithm cancels out the spikes: no shards are moved to address the write load hotspot.</p></li><li><p>Index-level hotspots can occur when a disproportionate number of index shards are assigned to the same node, rather than spreading out across nodes, despite the index balance weight in the linear algorithm. Index balance weight can, in some situations, be little compared to the other weight factors. It can also get skewed and counterbalanced by another non-average individual weight in the linear algorithm, as described in a previous bullet.</p></li><li><p>No search load balancing.</p></li><li><p>Regular indices do not have write load estimate support, leaving some write load hotspots unaddressed. Only data stream indices have write load estimates.</p></li><li><p>Write load hotspots can be missed. Write load estimates are only refreshed at rollover time, which can be infrequent in some configurations, causing new load to be ignored for some time. The write load is also the average write load activity over a potentially large window of time between index rollover events, so temporary write load increases can disappear when averaged with inactive write periods.</p></li></ul><p>The above issues persist in some Elasticsearch deployments and require monitoring and workload tuning to manage when they do occur. Shard allocation balancing in Elasticsearch Serverless aims to address these issues and avoid any manual intervention requirements using a new approach that is explained in subsequent sections of this article.</p><h2>Elasticsearch Serverless Shard Allocation </h2><p>Elasticsearch Serverless considers node resources individually: shards are rebalanced away from a node when any resource usage on that node grows to threaten performance, and shard movements to a node are declined when the assignment could threaten that node’s performance.</p><p>The Elasticsearch single combined score per node is replaced in Elasticsearch Serverless with independent per-resource decisions:</p><p>
</p><p><strong>Elasticsearch Weights-Based Balancing</strong></p><p><strong>Elasticsearch Serverless Resource-Aware Deciders</strong></p><p><strong>Decision Basis</strong></p><p>Single weighted sum across four metrics</p><p>Each resource evaluated independently</p><p><strong>Metric Interaction</strong></p><p>A high value can offset a low one</p><p>No offsetting; each decider acts separately</p><p><strong>Decision Types</strong></p><p><code>YES</code> / <code>NO</code></p><p><code>YES</code> / <code>NO</code> / <code>NOT_PREFERRED</code></p><p><strong>Rebalancing Trigger</strong></p><p>Weight delta across nodes exceeds <code>threshold</code></p><p>Individually configurable safe limits per resource</p><p><strong>Explainability</strong></p><p>Can only make an educated guess</p><p>Each move traces to a named decider</p><p>The Elasticsearch Balancer has three phases, in order of priority, for shard movement decisions. The first phase is to assign unassigned shards. Assignment of unassigned shards is the top priority for data availability reasons. The second phase is to move shards that can no longer remain where they are assigned due to cluster configuration changes. Internally, <code>AllocationDecider</code> implementations enforce cluster settings, like <a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/shard-allocation#index-allocation-filters">index-level shard allocation filtering</a>, <a href="https://www.elastic.co/docs/deploy-manage/distributed-architecture/shard-allocation-relocation-recovery/shard-allocation-awareness">shard allocation awareness</a>, <a href="https://www.elastic.co/docs/reference/elasticsearch/configuration-reference/cluster-level-shard-allocation-routing-settings#disk-based-shard-allocation">disk usage thresholds</a>, or moving shards off of a node before shutdown. The third phase rebalances shards when the <code>cluster.routing.allocation.balance.threshold</code> is exceeded, using the previously described weights algorithm.</p><p>The new Serverless balancing approach adds additional logic to the Balancer’s first and second phases, leveraging the existing <code>AllocationDecider</code> logic, and eliminates the third phase. Previously, each <code>AllocationDecider</code> had simple responses of <code>YES</code> and <code>NO</code>. Now, the decision type of <code>NOT_PREFERRED</code> has been added, along with several new <code>AllocationDecider</code> implementations. An <code>AllocationDecider</code> will return <code>NOT_PREFERRED</code> when it observes that performance might suffer from a shard’s assignment to a particular cluster node. The Serverless Balancer will prefer a node assignment for the shard where all <code>AllocationDecider</code> implementations reply <code>YES</code>.</p><p>A <code>NOT_PREFERRED</code> shard allocation may be left uncorrected if all other node assignments return <code>NO</code> or <code>NOT_PREFERRED</code>. Such responses mean that the shard cannot be assigned elsewhere without either violating a cluster/index rule or potentially degrading the performance of another cluster node. Serverless Autoscaling activates before all cluster nodes hotspot: even one unaddressable hotspot leads to a scale-up event. New <code>AllocationDecider</code> implementations have also been added for important finite resources, like available heap memory (further discussion below), using only the original <code>YES</code> and <code>NO</code> decisions: exceeding certain categories of resources can lead to node unavailability.</p><p>The individual weight metrics in the Balancer’s linear algorithm have been replaced by resource-aware <code>AllocationDecider</code> implementations, and new <code>AllocationDecider</code> implementations are being built for additional resources: Serverless Search Tier load-balancing improvements are currently in development. Each shard migration will have a clear purpose to address a potential resource usage bottleneck.</p><p>Internal stats have shown far fewer shard movements in general, without any noticeable accompanying node performance degradations – one workload showed a 50% reduction in shard movements with the same write throughput. Fewer shard movements has the benefit of: avoiding momentary read/write latencies from warming up local caches; and saving on cloud infrastructure costs moving data between servers.</p><h3>Serverless IndexBalanceDecider: Avoid Colocation of Index Shards</h3><p>The <code>IndexBalanceDecider</code> ensures index shard anti-affinity much more strictly than the original weights-based linear algorithm could achieve. Colocation of index shards in excess of the index’s average shards per available node is avoided, except in the case of a strict <code>NO</code> assignment (essentially non-existent right now in Serverless except for shutting down nodes and rolling upgrade incompatible version checks) or <code>NOT_PREFERRED</code> assignment due to temporary node hotspotting.</p><p>The <code>IndexBalanceDecider</code> is a very effective means of pre-balancing both write load and search load before user workloads begin to generate load statistics: each index begins life with its shards distributed across as many nodes as possible.</p><h4>IndexBalanceDecider Results: Even Write Load Distribution Across Cluster Nodes</h4><p>Write load across data nodes became much more evenly distributed after the <code>IndexBalanceDecider</code> was enabled in the Serverless Production environment. Projects fleet-wide generally show even ingest load (counted in saturated <code>WRITE</code> threadpool threads), combining the release of the <code>IndexBalanceDecider</code> and many other prior improvements:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb1d22b863724026b/6a969c19ecdaa7898505223b/4.png" alt="Ingestion load per data node across an Elasticsearch Serverless project, showing even write load distribution over time" /><p>A reproducible workload demonstrates a clear before and after view of the impact of the new <code>IndexBalanceDecider</code> when an ingest workload was run with and without it enabled:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb3788a358c3a3043/6a969c385c3126893d43ef2b/5.png" alt="Scale test ingestion load per node with the IndexBalanceDecider off then on, showing write load spread across data nodes" /><p>The <code>IndexBalanceDecider</code> also serves in the Serverless Search Tier to distribute shards of the same index as much as allowed, similarly limited only by the tier’s node count and the number of shards in each index.</p><h3>Serverless HeapUsageDecider: Assign Shards by Available Heap</h3><p>The <code>HeapUsageDecider</code> limits shard count on a node based on available heap to hold in-memory shard metadata and run associated write/read operations, removing the dependency on shard count limits per node. The <code>HeapUsageDecider</code> returns a strict <code>YES</code> or <code>NO</code> decision, rather than using the new <code>NOT_PREFERRED</code> decision type, because a data node risks an OOM event if the estimated available heap memory is exceeded.</p><h4>HeapUsageDecider Results: Reduced Data Node OOMs</h4><p>Data node OOMs in the serverless index tier decreased significantly as the <code>HeapUsageDecider</code> rolled out to the Serverless production environment.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5c28891290e3a635/6a969c88d04dac5ed56ca2f1/6.png" alt="Indexing OOM errors on Elasticsearch Serverless data nodes falling to near zero after the HeapUsageDecider rollout" /><p>Index tier OOM errors still occur from time to time in the Serverless Index Tier, though at a much reduced rate, as miscellaneous runaway memory usage edge cases are surfaced. The remaining OOM errors are being progressively resolved as they are identified, through a combination of memory usage improvements in the code, adding component level limits, and updating the internal Elasticsearch Serverless memory model service to more completely account for memory usage.</p><p>The <code>HeapUsageDecider</code> is not yet turned on in the Serverless Search Tier, due to the need for additional and different metrics, but that work is in active development.</p><h3>Serverless WriteLoadDecider: Prevent and Correct Write Load Hotspots</h3><p>The <code>WriteLoadDecider</code> receives periodically refreshed (every 30 seconds by default) per shard and per node write load stats and uses the data to correct and avoid write load hotspots. The master node retrieves stats directly from each data node’s write threadpool: an Elasticsearch node tracks the total time that its <code>WRITE</code> threadpools is in use, and each individual Elasticsearch shard instance tracks how much time it spent using its node’s <code>WRITE</code> threadpool.</p><p>A write load hotspot is identified at the node level. The criteria for a hotspot is the presence of <code>WRITE</code> threadpool queue latency above a configured threshold and sufficiently high, and sustained, <code>WRITE</code> threadpool thread saturation. Once that situation is detected, the Balancer is signaled to select shards to move away from a hotspotting node, until fresh non-hotspotting write load stats are received from the node. The Balancer will do nothing if all nodes are hotspotting at once, expecting the Autoscaler to solve the problem by introducing more, or bigger, data nodes to the cluster.</p><p>The <code>WriteLoadDecider</code> uses a heuristic to choose shards to move away from a hotspotting node that aims to minimize ingest disruptions while still effectively reducing a node’s write load. A shard write load <code>threshold</code> is identified on a hotspotting node: the <code>threshold</code> is currently calculated as ½ the ingest load of the hottest shard on that node. Shards that can be moved are then prioritized in the following order:</p><p><code>threshold</code><code> = ½ * </code><code>maxWriteLoadShardOnNode</code></p><ol><li><p>Shards with write load in the range [<code>threshold</code>, <code>maxWriteLoadShardOnNode</code>), the shard at or closest to threshold preferred.</p></li><li><p>Shards with write load in the range (<code>threshold</code>, <code>0</code>], the shard closest to threshold preferred.</p></li><li><p>Shards with write load equal to <code>maxWriteLoadShardOnNode</code>.</p></li><li><p>Shards with zero write load.</p></li></ol><p>The heuristic prefers to avoid disruption to the highest ingest shards and instead chooses middlingly loaded shards. Movement of the hottest shard will cause the most latency disruption; and movement of the coldest shards will be the least effective in resolving a hotspot.</p><p>The Balancer limits write load hotspot correction shard moves to one move per hotspotting node per stats refresh period, in order to see the effect of a move in real-time node-level write load, before attempting any further corrections. This was a simple initial design that proved effective. Furthermore, the Balancer will not move a shard whose write load alone is sufficient to meet the node-level hotspot criteria: this would just relocate a hotspot to another data node, not actually resolve the hotspot. The Serverless Autoscaler and Serverless Autosharding components are relied upon to resolve hotspots that reallocation of shards cannot.</p><p>The <code>WriteLoadDecider</code> returns <code>NOT_PREFERRED</code> when acceptance of a shard could cause a node to start experiencing <code>WRITE</code> threadpool queue latency and create a hotspot. A shard will still be relocated to a <code>NOT_PREFERRED</code> node, however, and risk some performance degradation, as a better option than, say, risking a data node OOM from keeping a shard on a data node where the <code>HeapUsageDecider</code> returns <code>NO</code>.</p><h4>WriteLoadDecider Results: Hotspots are Quickly Corrected </h4><p>Hotspot stats showed general improvement as the <code>WriteLoadDecider</code> was rolled out to Serverless production, in particular the fleet-wide time to correct a hotspot decreased greatly:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1e3ad68727d2121d/6a969cb28814aa3f7c89c987/7.png" alt="Write load hotspot duration p50, p95 and p100 in Elasticsearch Serverless dropping after the WriteLoadDecider rollout" /><p>Since these graphs were collected, additional work has been released incrementally to better prevent and correct hotspots, and improvements are still in progress.</p><h2>Serverless Autoscaling, Autobalancing, and Autosharding</h2><p>Elasticsearch Serverless relies on both new autobalancing logic and new autoscaling logic. The Serverless Balancer must sufficiently distribute shard resource usage across nodes in order to fully saturate the cluster’s resources. The Serverless Autoscaler will trigger a scale-up event when it receives a report that a certain percentage of the total cluster resources are in use and more resources are needed. The Autoscaler will not scale up the cluster if one node is hotspotting and another node has an excess of available resources because the resources are summed across nodes. Therefore, the Balancer must first do a good job on load distribution, and then the Autoscaler will activate as needed.</p><p>Autosharding based on write load is also in progress and coming soon to address shard hotspots. Elasticsearch Serverless projects have a default number of shards per index based on the project type. These defaults generally work, but do not account for all possible workloads. Hotspots can occur when an index has too few shards, as well as too many. Too few index shards leads to the Balancer being unable to further distribute an index’s write load across available data nodes, and then the Autoscaler will not see a problem because the cluster-level resources are not fully consumed. Conversely, indices cannot by default have too many shards, since that could degrade search performance for small indices and potentially strain cluster metadata operations if the total number of shards in a cluster grew too large.</p><h2>Production Example: 708 TB Data Set, 37 Index Tier Nodes (not counting Search Tier), 4,100 Indices, 30,000 Shards</h2><p>The following graphs cover a period when the Index Tier, in an Elasticsearch Serverless project, scales up from 10 to 37 indexing nodes and then back down to 10 after a write load spike dissipated.</p><h3>Graph of the Ingest Load Per Index Node</h3><p>This graph shows fairly even distribution of load, though a little less even temporarily during scale-up. There are nearly 250 fully saturated write threads at peak load. </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6daf46a0823c4a37/6a969cd437d7f3a2008e9204/8.png" alt="Ingestion load per Elasticsearch Serverless index node, peaking near 250 saturated write threads during a load spike" /><h3>Graph of CPU Saturation Per Index Node</h3><p>CPU usage remains within safe bounds. Usage is mostly below 60%, except for momentary outliers that reach into the 90% range as write load rises before nodes are added to the cluster.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf60cff9da89290c9/6a969d203eabd090dd4409b4/9.png" alt="CPU usage per Elasticsearch Serverless data node, mostly below 60% with brief peaks above 90% during scale-up" /><h3>Graph of WRITE Threadpool Queue Latency Per Node </h3><p>When a node’s <code>WRITE</code> threadpool is fully saturated, tasks are placed in the threadpool’s queue. Queuing can happen with few tasks, if active write tasks are long-running, or there may simply be a lot of tasks.</p><p>This graph’s time window is zoomed in further than the others. One node reaches 75 seconds of queue latency during the scale-up spike. There are 29 nodes when the queue latency spike occurs at 19h25m, before autoscaling calls for 37 nodes at 19h28m.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9c16dbf83897366c/6a969d478814aa399089c98b/10.png" alt="Maximum WRITE threadpool queue latency per data node, with one Elasticsearch Serverless node reaching 75 seconds" /><h3>Graphs of Total Cluster Ingest per Second, in Documents and MBs</h3><p>Ingest rate peaks at 190,000 documents / second and 54.40MB / second. The document ingest rate is respectable at 4000-5000 docs/sec per node. The MBs ingest rate, however, is very low in this case: this can happen when indexing operations involve heavy computation. Document ingestion rate can also vary depending on the size of the documents.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8cec5b62e8339a7f/6a969d7e37d7f35c438e9208/12.png" alt="Total indexing request rate for an Elasticsearch Serverless cluster, peaking at 190,000 documents per second" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta9c147b7bd608b73/6a969dae5f9db76cd1560dc8/11.png" alt="Bulk byte indexing rate for an Elasticsearch Serverless cluster, peaking at 54.40 MB per second during the write spike" /><h2>What’s Next for Elasticsearch Serverless Balancing</h2><p>The team is currently working on shard balancing improvements for the Serverless Search Tier, focusing on creating metrics and <code>AllocationDecider</code> implementations for search performance. The team is excited to share these improvements soon!</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-shard-balancing-serverless</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-shard-balancing-serverless</guid>
    <category><![CDATA[Elastic Cloud Serverless]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Dianna Hohensee]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltca84dead3c455000/6a969ba12707c589983290d7/1.png" length="0" type="image/png"/>
    <pubDate>Tue, 01 Sep 2026 15:25:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Migrating 1,100 files to Redux Toolkit v2 without freezing the Kibana monorepo]]></title>
    <description><![CDATA[Kibana gave Redux Toolkit v2 the default package name and pushed v1 onto an explicit alias, which inverts the usual migration order. Webpack externals, yarn resolutions and an ESLint rule keep React Redux v7 and v9 out of each other's way.]]></description>
    <content:encoded><![CDATA[<p>We moved roughly 1,100 files in the Kibana monorepo onto <a href="https://github.com/elastic/kibana/pull/235577">Redux Toolkit (RTK) v2</a> aliases without asking a single plugin team to pause feature work. The usual migration pattern runs the other way around. Default package names (<code>@reduxjs/toolkit</code>, <code>react-redux</code>, <code>redux</code>) now resolve to v2, and existing v1 code sits behind explicit aliases, like <code>redux-toolkit-v1</code> and <code>react-redux-v7</code>. Both versions live in <code>node_modules</code> at once, kept apart at runtime by npm aliases and webpack module replacement. An ESLint rule scoped to 36 plugin paths catches anything that tries to cross. When a team is ready, it deletes its path from that list and switches back to the default imports, and the teams around it carry on shipping.</p><h2>Why upgrade to Redux Toolkit v2?</h2><p>RTK v2 was released in late 2023. That's nearly three years of running on a major version behind in one of the most widely used state management libraries in the JavaScript ecosystem. It reflects how hard this upgrade is in a codebase of Kibana's size. A <a href="https://github.com/elastic/kibana/pull/178986">previous attempt</a> tried the big-bang approach and stalled when the real scope became clearer. </p><p>So what does v2 actually bring? It ships alongside Redux core 5.0, React-Redux 9.0, Reselect 5.0, and Redux Thunk 3.0. React-Redux 9.0 requires React 18 and drops the <code>useSyncExternalStore</code> shim that v8 carried for React 16/17. Since Kibana already runs React 18, upgrading sheds legacy compatibility code and keeps Kibana on the actively maintained Redux majors.</p><p>RTK v2 also brings genuinely useful new features, including inline selectors in <code>createSlice</code> and opt-in inline async thunks through a customized <code>buildCreateSlice</code> setup, along with a <code>combineSlices</code> API with slice reducer injection for code splitting. That last one is particularly interesting for Kibana's plugin architecture where lazy-loading is the norm.</p><h2>How Redux is used across the Kibana monorepo</h2><p>Before diving into the solution, it's worth understanding just how varied Redux usage is across Kibana. A full audit of the codebase (tracked in <a href="https://github.com/elastic/kibana/issues/239863">#239863</a>) revealed several distinct camps:</p><p><strong>Pattern</strong></p><p><strong>Plugins and packages</strong></p><p><strong>What the migration needs</strong></p><p>Redux Toolkit v1</p><p>Discover, Lens, Synthetics, Security Solution</p><p>Full v1 to v2 migration</p><p>Plain Redux v4</p><p>Canvas, Maps, Index Management, Cross-Cluster Replication</p><p><code>redux-v4</code> alias only, no RTK migration</p><p>Kea</p><p>Enterprise Search (150+ files), Content Connectors</p><p><code>react-redux-v7</code> alias, no RTK migration</p><p><code>redux-saga</code></p><p>Synthetics, Graph, Uptime</p><p>Store setup only, saga is version-independent</p><p><code>typescript-fsa</code></p><p>Security Solution data-table package</p><p>Out of scope</p><p>Types and single imports</p><p>Expressions, Monitoring</p><p>Alias swap only</p><p>Plugins such as Discover, Lens, Synthetics, and Security Solution use RTK v1 APIs,including <code>createSlice</code>, <code>configureStore</code>, <code>createAsyncThunk</code>, and <code>createSelector</code>. These are the ones that actually need the v1 to v2 migration. But even here, complexity varies wildly. Lens uses stand-alone <code>getDefaultMiddleware</code> (removed in v2) and <code>PreloadedState</code> (also removed). Security Solution is the largest consumer at 300+ files, mixing modern RTK with legacy plain Redux patterns.</p><p>Canvas, Maps, Index Management, Cross-Cluster Replication, and several others still use plain Redux v4via <code>createStore</code>, <code>combineReducers</code>, <code>applyMiddleware</code>, and <code>connect</code>, which are classic patterns from the pre-RTK era. These don't need RTK migration at all since they're not using it  in the first place, but they do need the <code>redux-v4</code> alias since the default <code>redux</code> package is now v5.</p><p>Enterprise Search and Content Connectors use <code>kea</code>, a Redux abstraction layer with its own logic builders (<code>kea()</code>, <code>useValues</code>, <code>useActions</code>). There are more than 150 files in Enterprise Search alone. RTK migration isn't applicable here, since <code>kea</code> is its own world. But it <em>does</em> depend on <code>react-redux</code> v7 under the hood, which is where the bundler tricks come in.</p><p>Synthetics, Graph, and Uptime use <code>redux-saga</code> for side effects. Saga integration is actually independent of the RTK version, but these plugins need their store setup migrated.</p><p>The Security Solution data-table package uses <code>typescript-fsa</code> and <code>typescript-fsa-reducers</code> instead of RTK entirely, with its reducer embedded into Security Solution's main store, and isn’t part of the RTK migration at all.</p><p>The Expressions plugin only imports <code>shallowEqual</code> from <code>react-redux</code>, and Monitoring only imports types. These just need an alias swap.</p><p>Asking every team to migrate simultaneously was a nonstarter. The breaking changes in RTK v2 include stricter type checking and removed APIs, like <code>enableES5()</code> from immer, <code>getDefaultMiddleware</code> and <code>PreloadedState</code> gone entirely, <code>AnyAction</code> replaced by <code>UnknownAction</code>, and behavioral changes in how middleware is configured.</p><h2>Running Redux Toolkit v1 and v2 side by side</h2><p>The solution was to flip the typical migration pattern on its head. Instead of keeping the default imports on v1 and introducing v2 under aliases, the default package names (for example, <code>@reduxjs/toolkit</code>, <code>react-redux</code>, and <code>redux</code>) now point to v2. The old versions live under versioned aliases:</p><ul><li><p><code>redux-toolkit-v1</code></p></li><li><p><code>react-redux-v7</code></p></li><li><p><code>redux-v4</code></p></li><li><p><code>immer-v9</code></p></li><li><p><code>reselect-v4</code></p></li><li><p><code>redux-thunk-v2</code></p></li></ul>{
"@reduxjs/toolkit": "2.12.0",
"redux-toolkit-v1": "npm:@reduxjs/toolkit@1.9.7",
"react-redux": "9.2.0",
"react-redux-v7": "npm:react-redux@7.2.8"
}<p>This is npm's alias syntax. <code>"react-redux-v7": "npm:react-redux@7.2.8"</code> installs the old version under a different name. Both versions coexist in <code>node_modules</code> without conflicts.</p><p>The insight here is that all existing code in this pull request (PR) was moved to v1 aliases. Every <code>import { useSelector } from 'react-redux'</code> became <code>import { useSelector } from 'react-redux-v7'</code>. That's ~1,100 files touched, but the vast majority (~1,000) are mechanical one-liner import swaps. When a team is ready to migrate to v2, they switch back to the default import names. Once all v1 aliases disappear from the codebase, the old packages can be removed entirely.</p><p>This avoids the alternative, where v2 imports would end up under nonstandard names permanently, leaving nonstandard imports in the codebase for the long term.</p><h2>Serving both versions through the bundler</h2><p>Getting two versions of the same library to coexist at runtime is where things got interesting. Kibana uses <code>kbn-ui-shared-deps-npm</code> to bundle common dependencies as shared webpack externals. This needed to serve both the new v2 packages <em>and</em> the v1 aliases so that both are available at runtime.</p><h3>Pinning @elastic/charts with yarn resolutions</h3><p>Then there's <code>@elastic/charts</code>. It depends on RTK v1 internally and can't just be upgraded independently since it's an upstream package. Yarn resolutions pin its nested dependencies to v1 versions:</p>{
"@elastic/charts/@reduxjs/toolkit": "npm:@reduxjs/toolkit@1.9.7"
}<p>A <code>NormalModuleReplacementPlugin</code> in the shared deps webpack config detects when an import of <code>immer</code>, <code>@reduxjs/toolkit</code>, <code>redux</code>, <code>react-redux</code>, or <code>reselect</code> originates from within <code>@elastic/charts</code> and redirects resolution to the nested v1 copies. This ensures that <code>@elastic/charts</code> resolves to its compatible v1 dependency set.</p><h3>Keeping Kea on React Redux v7 with webpack externals</h3><p>The <code>kea</code> library was another fun case. It declares <code>react-redux</code> as a peer dependency (<code>&gt;= 7</code>), so without special handling its imports resolve to Kibana's default v9 package. The migration keeps Kea consumers on <code>react-redux-v7</code>, so Kea must use that same React context. The fix uses function-based webpack/rspack externals that skip externalizing <code>react-redux</code> when the import comes from <code>node_modules/kea</code>, combined with a <code>NormalModuleReplacementPlugin</code> that rewrites it to <code>react-redux-v7</code>. This ensures that <code>kea</code> uses the v7 React context that matches the <code>&lt;Provider&gt;</code> wrapping its consumers.</p><p>Both the webpack (<code>kbn-optimizer</code>) and the rspack (<code>kbn-rspack-optimizer</code>) configs needed these changes, with a shared <code>isKeaReactReduxImport</code> helper extracted to keep the logic consistent.</p><h2>Using an ESLint rule to prevent cross-version imports</h2><p>With two versions available, accidental cross-version imports are the biggest risk. A new <code>@kbn/imports/no_redux_toolkit_v2_imports</code> ESLint rule catches any import of the v2 default packages (such as <code>@reduxjs/toolkit</code>, <code>react-redux</code>, or <code>redux</code>, among others) in code that hasn't been migrated yet. It even auto-fixes them to the v1 aliases for file imports and Jest mock paths.</p><p>The rule is scoped via an override in <code>.eslintrc.js</code> to the ~36 plugin and package paths currently using v1. When a team migrates, they simply remove their path from the override list. This clean, self-service approach requires no coordination.</p>// .eslintrc.js (simplified)
overrides: [{
  files: [
'src/platform/plugins/shared/discover/**/*.{ts,tsx}',
'src/platform/plugins/shared/workflows_management/**/*.{ts,tsx}',
// ... 34 more paths
],
  rules: {
'@kbn/imports/no_redux_toolkit_v2_imports': 'error',
  },
}]<h2>Why mixing React Redux v7 and v9 breaks the context</h2><p>This is worth calling out because it's an easy failure mode to miss during an upgrade. <code>react-redux</code> v9 and v7 create separate React contexts. If a component tree has a v9 <code>&lt;Provider&gt;</code> at the top but a child component calls <code>useSelector</code> from v7 (or vice versa), React-Redux cannot find the matching context. In development, it throws an error explaining that the component must be wrapped in a matching <code>&lt;Provider&gt;</code>; in production, the missing context causes a runtime error when the hook accesses the store.</p>Error: could not find react-redux context value; please ensure the component is wrapped in a &lt;Provider&gt;<p>This means that each plugin needs to be explicitly pinned to one version. Shared packages that use <code>react-redux</code> can only be consumed by code on the same version, since mixing isn't possible. This is a constraint that makes the migration inherently per plugin rather than per file.</p><h2>Migration batches: What can move independently</h2><p>The dual-version setup gives every team a clear path forward, and the dependency graph analysis from the tracking issue identified natural migration batches:</p><ul><li><p><strong>Batch 1: Independent, self-contained stores.</strong> Packages like <code>kbn-coloring</code>, <code>transform</code>, <code>timelines</code>, and <code>expandable-flyout</code> have fully internal Redux stores with no types leaking through their public APIs. These can be migrated independently by their owning teams, with minimal risk.</p></li></ul><ul><li><p><strong>Batch 2: Coupled packages.</strong> Some packages share RTK types across boundaries and <em>must</em> migrate together. The machine learning (ML)/artificial intelligence for IT operations (AIOps) chain is one example: <code>@kbn/ml-response-stream</code> exports a <code>streamSlice</code> (a <code>createSlice</code> return value) that <code>@kbn/aiops-log-rate-analysis</code> embeds directly into its <code>configureStore</code>. Migrating one without the other causes type mismatches between v1 and v2 slice types. Similar coupling exists across the Lens ecosystem. The Lens plugin depends on <code>@kbn/coloring</code> (which has its own RTK store), <code>@kbn/lens-embeddable-utils</code>, and <code>@kbn/lens-common</code>, while itself being consumed by 40+ packages and plugins across chart expressions, visualizations, Maps, Canvas, and observability plugins. Whether Redux types leak through a package's public API determines if it can be migrated independently or needs coordination. <code>kbn-coloring</code>'s store is internal to its React components so it's safe to migrate alone, but other coupling points need careful analysis.</p></li></ul><ul><li><p><strong>Batch 3+: The big ones.</strong> Discover, Security Solution, and Lens each have their own migration timelines. Security Solution's 300+ files and mix of RTK with plain Redux v4 and <code>typescript-fsa</code> make it the largest effort, but the different patterns can be addressed independently. Lens has the trickiest v2 breaking changes around middleware configuration; stand-alone <code>getDefaultMiddleware</code> and <code>PreloadedState</code> are both removed in v2, and it has four custom middleware files with complex typing.</p></li></ul><p>Beyond the batched migrations:</p><ul><li><p><strong>Deprecated features</strong> can stay on v1 aliases. When the feature is removed, the v1 imports disappear through code deletion, without any migration work.</p></li><li><p><strong>Plain Redux v4 plugins</strong> (Canvas, Maps, and others) are entirely out of scope for RTK migration. They'd benefit from modernization, but that's a separate initiative.</p></li><li><p><strong>Kea plugins</strong> need <code>react-redux-v7</code> to <code>react-redux</code> alias updates eventually, but no RTK migration. The longer-term question (whether to keep Kea or migrate to RTK v2) is a separate decision.</p></li><li><p><strong>The dual-version approach</strong> adds measurable bundle overhead during the transition. It’s a trade-off but is acceptable for the migration period.</p></li></ul><h2>Lessons for other large monorepo upgrades</h2><p>The ESLint rule turned out to be the linchpin. Without automated enforcement, aliased imports would drift back to default names within weeks. With it, the migration state is visible in the paths listed in the override. As of the initial PR, zero files import from <code>@reduxjs/toolkit</code> v2. Every RTK usage goes through the <code>redux-toolkit-v1</code> alias. That's the starting line.</p><p>The preparation work also reached beyond import paths. Jest mocks referencing <code>react-redux</code> needed updating to <code>react-redux-v7</code>, as did Storybook previews, test helpers, and ambient type declarations. Multiple rounds of <code>node scripts/eslint_all_files --no-cache --fix</code> caught the mechanical cases; the remaining cases needed manual fixes.</p><p>If you're facing a similar major dependency upgrade in a large monorepo, the pattern of giving the new version the default name and the old version an explicit alias is worth considering. New code naturally uses the current version, while older usage stays visible and trackable until it reaches zero.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/redux-toolkit-v2-migration-kibana-monorepo</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/redux-toolkit-v2-migration-kibana-monorepo</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Developer Experience]]></category>
    <category><![CDATA[Kibana]]></category>
    <dc:creator><![CDATA[Walter Rafelsberger]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3fa3c79a7e246c89/6a9111235c3126655043df06/unnamed.png" length="0" type="image/png"/>
    <pubDate>Fri, 28 Aug 2026 15:20:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Taming PUNKs: How ES|QL queries Elasticsearch fields it was never told about]]></title>
    <description><![CDATA[In Elasticsearch 9.5, ES|QL can query unmapped fields. It reads them from _source or returns nulls, so a query keeps working when a field drops out of the mapping and you avoid a reindex that takes hours.]]></description>
    <content:encoded><![CDATA[<p>How do you make an analytical query engine use data that it cannot know exists? You “just” read the query, since everything that the user asks for is right there. Right?</p><p>In Elasticsearch 9.5, <a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language (ES|QL)</a> queries no longer fail when a field isn't in the mapping. The new <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-unmapped-fields"><code>unmapped_fields</code> setting</a> lets queries load values from <code>_source</code> or fill with <code>nulls</code>, so queries keep working even when a backing index changes and a field goes missing, and can use unmapped data without reindexing. Here’s how we built that: the design choices and the edge cases (including a class of fields we nicknamed PUNKs), along with the testing strategies that gave us the confidence to ship it in general availability (GA).</p><h2>Why ES|QL queries fail when a field is unmapped</h2><p>You built a visualization using an ES|QL query. You refined it, and the query grew. You’re at 15 chained commands and counting, but it does <em>just</em> the right thing. It works, and your dashboard is <em>useful</em>.</p><p>Your query uses an index from a remote cluster, say <code>my-remote:logs-foo</code>. But actually, <code>logs-foo</code> is an alias, and at some point, the remote cluster makes it point to a different backing index. The new index is missing a field that’s used in your query, and your query and visualization break.</p><p>Or maybe you have an already fairly large index, and while building ES|QL queries on top of it, you realize that you’d like to use a field in the indexed documents that unfortunately never made it into the index mapping. You could reindex the data, but that would take hours.</p><p>ES|QL’s <code>unmapped_fields</code> setting is meant to deal with these types of situations.</p><p>If your query looks like this:</p><p>and <code>some_field</code> is unmapped, ES|QL’s default behavior is to fail with a verification exception.</p><p>You can use the <code>unmapped_fields</code> setting to instead either fill <code>some_field</code> with <code>null</code>s or read it from the document’s <code>_source</code>, like so:</p><h2>How ES|QL resolves queries with field caps</h2><p>Before we jump into the inner workings of <code>unmapped_fields</code>, we have to look into how ES|QL resolves queries regularly. Let’s consider the above query:</p><p>We said that if <code>some_field</code> isn’t in the mapping for <code>index</code>, ES|QL will reject the query. How does it make that decision?</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7ba9af4b4b550018/6a8ee83fe41d7fea88654d42/image4.png" alt="ES|QL query resolution flow: analyzer checks index mappings, unresolved fields fail with Unknown column error" /><h3>How field caps tells ES|QL which fields exist</h3><p>In a typical schema-on-write fashion, Elasticsearch clusters maintain mappings with their respective indices. As a first step, ES|QL makes an internal request to the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-field-caps">field caps endpoint</a> to determine which fields the <code>index</code> has. It then passes the query, together with the field caps response, to the query planner, which consists essentially of the query analyzer (unrelated to analyzers of text fields) and query optimizer. The analyzer makes sense of raw names, like <code>some_field</code>, and notices that they correspond to index fields (or not). If all went well, the query is then passed on to the optimizer, which rewrites the query for efficiency, before it’s handed to the compute engine for execution.</p><h3>How the analyzer resolves field names in the query plan</h3><p>Let’s zoom in to the analyzer. The parsed query is represented in a tree structure, and the analyzer partially rewrites it, one command at a time, until it either has resolved all references or not.</p><p>For illustration, let’s use a somewhat more complex query and see how the analyzer would resolve it:</p><p>The parsed tree is actually a chain here, and it looks something like this:</p><p></p><p>The analyzer then moves up through the query tree to try and resolve the field names used in every command.</p><p>This is a simplified version of how we represent parse trees in tests and when debugging. The bottom of the chain corresponds to the <code>FROM</code> command and contains a list of all mapped fields that we know about, obtained from the field caps endpoint. (The <code>{f}</code> suffix marks an actually mapped field for better distinction later.)</p><p>The two <code>EVAL</code> nodes on top of it correspond to the remaining commands, and their fields are still unresolved, expressed by the question mark <code>?</code> in front of the name. At this point, the analyzer still has to check whether they correspond to existing index fields.</p><p>For the <code>EVAL</code> that defines <code>uppercased_mapped</code>, it can see that the previous command outputs <code>mapped_field</code>, so the unresolved <code>?mapped_field</code> marker can be replaced by a real field reference:</p><p>Next, it encounters the topmost <code>EVAL</code>, which defines <code>uppercased_unmapped</code>. The previous tree nodes produce only two fields: <code>[mapped_field, uppercased_mapped]</code>. The reference <code>?unmapped_field</code> thus has to remain unresolved. We bail here and emit the verification exception to the user.</p><h2>How unmapped_fields LOAD and NULLIFY work</h2><h3>Adding unmapped fields to the query plan</h3><p>When using <code>unmapped_fields=”NULLIFY”</code> or <code>”LOAD”</code>, we do something else; we act as if the field was actually in the index. The analyzer adds <code>unmapped_field</code> to the <code>From</code> node and marks it as unmapped to signal to the compute engine that this has to be read from <code>_source</code> or filled with <code>null</code>s. Let’s express this with a <code>{u}</code> (for <strong>u</strong>nmapped):</p><p>After amending the <code>From</code>, the analyzer can continue trying to resolve the topmost <code>Eval</code> node. It sees that the upstream nodes produce the fields <code>[mapped_field, unmapped_field, uppercased_mapped]</code> and thus <code>unmapped_field</code> can be correctly resolved:</p><p></p><p>The query plan is now fully resolved and can be passed down the regular optimization-execution pipeline. Other than the actual value extraction mechanism, everything stays the same. Schematically, the workflow looks like this:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt59bf2d2da201d00a/6a8ee8df8658b77c28469356/image2.png" alt="ES|QL unmapped fields flow: analyzer retries with NULLIFY or LOAD instead of failing on an unresolved field" /><h3>Example: enabling unmapped fields with the SET directive</h3><p>To give an example, let’s fire up a cluster and create an index with non-dynamic mappings.</p>PUT /index
{                                 
  "mappings": {
    "dynamic": false,
    "properties": {
      "mapped_field": {"type": "keyword"}
    }
  }
}

POST /index/_doc?refresh
{
  "mapped_field":"foo"
  "unmapped_field": "bar"
}<p>We can run the example query, above:</p>POST /_query
{
  "query": """
           FROM index
           | EVAL uppercased_mapped = TO_UPPER(mapped_field)
           | EVAL uppercased_unmapped = TO_UPPER(unmapped_field)
           """
}<p>This should result in the error message:</p><p><code>Unknown column [unmapped_field], did you mean [mapped_field]?</code></p><p>To make things work, we can prepend <code>SET unmapped_fields=”...”;</code> with <code>LOAD</code> or <code>NULLIFY</code>:</p>POST /_query
{
  "query": """
           SET unmapped_fields="LOAD";
           FROM index
           | EVAL uppercased_mapped = TO_UPPER(mapped_field)
           | EVAL uppercased_unmapped = TO_UPPER(unmapped_field)
           """
}

 mapped_field  |unmapped_field |uppercased_mapped|uppercased_unmapped
---------------+---------------+-----------------+-------------------
foo            |bar            |FOO              |BAR<h3>Inspecting the analyzer's rewrite steps</h3><p>If you want to see what the query analyzer is doing to the parse tree, you can log the query rewrite steps, like so:</p>PUT /_cluster/settings"
{
  "transient" : {
    "logger.org.elasticsearch.xpack.esql.analysis.Analyzer.changes": "TRACE"
  }
}<p>This will log a line containing <code>Rule rules.ResolveUnmapped applied with change…</code> You’ll see that <code>unmapped_field</code> is added to the bottom of the parse tree as described above.</p><h2>Why we have to infer the schema</h2><p>Of course, this isn’t the only possible method to deal with unmapped fields. Here are some alternatives:</p><ol><li><p>We could also scan or probe the documents in <code>index</code> to determine that their <code>_source</code> actually has the <code>unmapped_field</code>.</p></li><li><p>We could disable verifications in the analyzer and make the compute engine blindly pass unmapped fields through individual computation steps.</p></li></ol><p>The first alternative front-loads more work to understand the <em>actual</em> schema of an index and thus generally increases latency. It doesn’t scale to large, highly distributed datasets. The second alternative isn’t viable since it means a large-scale change to how ES|QL’s compute engine is built, because it passes around streams of data with fixed columns from one operator to another.</p><p>In contrast, the approach we chose is neatly compatible with ES|QL’s existing optimization pipeline.</p><p>The trade-off is that the analyzer has to correctly <em>infer</em> a schema based on the actual index mappings (obtained from the field caps endpoint) and additional fields used inside the query. </p><p>This isn’t always straightforward. There were two main challenges:</p><ol><li><p>There are many different query shapes and commands that can be used. The mechanism needs to detect unmapped fields, update the proper <code>FROM</code> command, and pass the new field through the halfway resolved plan correctly in all cases.</p></li><li><p>There are many different mappings we have to deal with, and we specifically need to make our feature work correctly when mappings <em>change over time</em> on top of that.</p></li></ol><p>In the following, we’ll focus on <code>LOAD</code>, although some problems (generally many fewer) also apply to <code>NULLIFY</code>.</p><h3>Which index to load unmapped fields from for LOOKUP JOIN and FORK</h3><p>To briefly illustrate the first problem, here are some choices we needed to make:</p><ul><li><p>Which index do we load from when using <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/lookup-join">lookup joins</a>? This one?</p><p>The <code>unmapped_field</code> cannot be attributed to both indices. We chose <code>index</code> since this is where we expect mappings to change more often than in lookup indices.</p></li><li><p>Similarly, how do we deal with subqueries and views or the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/fork"><code>FORK</code> command</a>? In the following:</p><p>one fork branch triggers loading of an unmapped field. Is it also present in the other fork branch? (Yes, it should be, but it’s not obvious and is specifically not true if the two <code>FORK</code>s are replaced by independent subqueries.)</p></li></ul><h3>Two principles to keep queries working</h3><p>The second problem, diversity of mappings and their evolution over time, is a far bigger driver for complexity. We strove for two basic usability principles:</p><ul><li><p>Queries that work in the default mode should generally still work when using <code>unmapped_fields=”NULLIFY”</code> and <code>”LOAD”</code>.</p></li></ul><ul><li><p>Queries that work when all fields are mapped should generally still work with <code>NULLIFY</code> and <code>LOAD</code> when a field becomes unmapped and vice versa.</p></li></ul><h3>The type of unmapped fields and inadvertent type conflicts</h3><p>Let’s talk about data types to see where this leads to complexity. First, when using <code>unmapped_fields=”LOAD”</code>, we need to assume a data type for unmapped fields. We chose <code>KEYWORD</code>, which allows us to avoid type conflicts when reading from <code>_source</code>. One document can contain <code>”unmapped_field”: “foo”</code>, and another can contain <code>”unmapped_field”: 123.4</code>. It’s fine because we treat both as strings.</p><p>However, this is a violation of the second principle when a non-<code>KEYWORD</code> field happens to go unmapped. Consider this query:</p><p>If <code>some_field</code> becomes unmapped, we’ll have to assume that the <code>KEYWORD</code> type and the query will fail with a type conflict.</p><p>Type conflicts <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-multi-index#esql-multi-index-invalid-mapping">aren’t new</a> and can be dealt with by using explicit casts in the query, like so:</p><p>It would be great if ES|QL just inferred a useful type to cast to, but this is something for the future.</p><h3>Type conflicts with partially unmapped fields, or: making PUNKs well behaved</h3><p>In addition to fully unmapped fields, <em>partially unmapped</em> fields are everywhere and should also work with <code>LOAD</code>. Let’s look at a query that uses multiple indices.</p><p>Let’s say that there are indices <code>index</code> and <code>index_without_some_field</code>, containing just the following documents.</p>// index1
{
  "some_field": "foo"
}

// index2
{
  "some_field": "bar"
}<p>Now let’s consider the query:</p>FROM index, index_without_some_field<p>and assume that <code>some_field</code> is unmapped in <code>index_without_some_field</code>. This will return:</p>some_field
-------------
 foo
 null<p>because ES|QL doesn’t load unmapped fields per default.</p><p>Of course, when setting <code>unmapped_fields=”LOAD”</code>, we want to load from <code>_source</code> for <code>index_without_some_field</code>:</p>SET unmapped_fields="LOAD";
FROM index, index_without_some_field

 some_field
-------------
 foo
 bar           // loaded from _source<p>As with fully unmapped fields, the case is simple when <code>some_field</code> is mapped as <code>KEYWORD</code> in <code>index</code>. When loading from <code>_source</code> for <code>index_without_some_field</code>,  we treat the field as <code>KEYWORD</code> as well, so there’s no conflict.</p><h3>What makes a field a PUNK</h3><p>The case is less clear when <code>some_field</code>is partially unmapped and the mapped leg is of a type other than <code>KEYWORD</code>. Such fields caused a lot of trouble until we found the best solution, which makes their acronym quite fitting: <strong>p</strong>artially <strong>u</strong>nmapped <strong>n</strong>on-<strong>k</strong>eyword fields, or PUNKs.</p><p>Unfortunately, PUNKs are far from being esoteric. For instance, it’s very natural to filter on a PUNK:</p><p>If <code>some_field</code> is mapped as <code>INTEGER</code> in <code>index</code>, the type conflict looks like this:</p><ul><li><p>Mapped as an <code>INTEGER</code> in <code>index</code>.</p></li><li><p>Unmapped in <code>index_without_some_field</code> and thus treated as <code>KEYWORD</code>.</p></li></ul><p>This can again be resolved manually by providing an explicit cast:</p><p>But this is far from acceptable. Even queries that work fine without <code>NULLIFY</code> and <code>LOAD</code> typically have <em>some</em> PUNKs; the unmapped leg is simply treated as <code>null</code> then. Both guiding principles are violated if <code>LOAD</code> requires an explicit cast here.</p><h3>Casting implicitly to the mapped type</h3><p>The solution is to introduce an implicit cast to the mapped type. In this case, we know that <code>some_field</code> is an <code>INTEGER</code> in <code>index</code>, and thus we treat it essentially as if the user wrote:</p><p>This means that queries that work without <code>LOAD</code> keep working. (ES|QL may even give you more data because we load the unmapped leg of PUNKs from <code>_source</code>.) Queries that used to work when a field is fully mapped also keep working when it goes unmapped in some (but not all) of its indices without having to alter the query in any way.</p><p><strong>Behavior</strong></p><p><strong>Default</strong></p><p><strong><code>NULLIFY</code></strong></p><p><strong><code>LOAD</code></strong></p><p>Unmapped field in query</p><p>Query fails</p><p>Query runs</p><p>Query runs</p><p>Values returned</p><p>None</p><p><code>null</code></p><p>Read from <code>_source</code> </p><p>Assumed type</p><p>n/a</p><p><code>NULL</code></p><p><code>KEYWORD</code></p><p>Partially unmapped field (PUNK)</p><p>Unmapped leg is <code>null</code></p><p>Unmapped leg is <code>null</code></p><p>Cast to the mapped type</p><p>Pushdown optimization</p><p>Full</p><p>Full</p><p>Per-node where fully mapped</p><h2>Don't throw it all away: Keeping ES|QL query optimization with unmapped fields</h2><p>There's one more thing to get right; that is, to make sure that optimizations still work correctly with <code>LOAD</code>. Consider the previous query:</p><p>ES|QL’s optimizer aggressively pushes down such <code>WHERE</code> filters and turns them into Lucene queries, so the compute engine doesn’t perform unnecessary work.</p><p>For this query, evaluating the filter in the compute engine would require fetching each and every document from the index; meaning, a full scan, very slow. If <code>some_field</code> was mapped as an <code>INTEGER</code> in both indices, we would instead perform a Lucene query, which looks like this:</p>{
  "range": {
    "some_field": {
      "gt" : 10,
      "boost" : 0.0
    }
  }
}<p>The compute engine then doesn’t have to load each document separately and check whether it matches the filter. Documents with <code>some_field &lt;= 10</code> are never fetched from the Lucene index, which is very efficient at this kind of filtering. Nice.</p><h3>Why filter pushdown is unsafe for unmapped fields</h3><p>If <code>some_field</code> is unmapped in <code>index_without_some_field</code>, however, it’s wrong to narrow the documents down using the same Lucene query, as Lucene interprets an unmapped <code>some_field</code> as <code>null</code> and thus no documents from <code>index_without_some_field</code> will ever match. This edge case is easy to miss, and it doesn’t help that there are several flavors of similar pushdowns. For instance, in the query:</p><p>the compute engine pushes even the counting to Lucene. Again, this is only correct if <code>some_field</code> is fully mapped.</p><p>This means that such optimizations can't apply to unmapped fields. It would be disappointing if a query used hundreds of indices and only one of them happened to not map <code>some_field</code>, causing the whole query to run unoptimized.</p><h3>How the local optimizer recovers the fast path</h3><p>Luckily, this problem has a solution, too. ES|QL actually has multiple optimizer runs:</p><ol><li><p>First, a preliminary optimizer run on the node handling the <code>_query</code> request.</p></li><li><p>Then, a second, local optimizer run on every node we fan out to because we need to fetch documents from its shards.</p></li></ol><p>The workflow after the initial optimization looks more like this:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb2c907ea31960f96/6a8eeabe36492afa7685fa36/image1.png" alt="" /><p>If the current node happens to map <code>some_field</code> in all shards, the local optimizer detects this situation and treats <code>some_field</code> like any other fully mapped field, including performing Lucene queries to greatly narrow down the dataset to be processed. In fact, data nodes process <code>LIMIT</code> queries like:</p><p>in batches of shards (to avoid loading too much data too eagerly), which includes a full local optimizer run per batch. This makes it even more likely to encounter batches where <code>some_field</code> is fully mapped, allowing ES|QL to run a fast Lucene query.</p><h2>Is it working now? Testing unmapped_fields across every ES|QL query shape</h2><p>As we have seen from the optimizer issues above, problems can hide in plain sight, even for very simple queries. Because <code>unmapped_fields=”LOAD”</code> can affect each and every kind of query, the surface area for bugs is essentially all of ES|QL.</p><p>Accordingly, getting good test coverage was tricky and challenged us to refine our testing strategies.</p><h3>Reusing spec tests with unmapped_fields</h3><p>Conveniently, ES|QL has an extensive corpus of test queries, together with expected result sets; we call them <em>spec tests</em> because they’re written using a simple text specification language, which looks roughly like this:</p>simpleEval
row a = 1 | eval b = 2
;

a:integer | b:integer
1         | 2
;<p>This lets us create new tests out of the existing ones by introducing slight variations. For instance, any existing test that runs without <code>SET unmapped_fields=”...”</code> should produce the exact same results when run with <code>SET unmapped_fields=”NULLIFY”</code>.</p><p>It also helped find major issues early in the development process, especially for <code>NULLIFY</code>. The <code>LOAD</code> setting changes the meaning of queries much more dramatically, limiting the usefulness of this approach. However, ES|QL also uses what we call <em>generative testing</em>; that is, we string together random commands, run the query, and then check whether the server reports a bug. This approach cannot confirm the correctness of results, but it still helped greatly with finding query types that didn’t work properly and resulted in some kind of error. (Property-based tests would be a refinement in the future by running the queries against a reference implementation. This way, correctness of results can also be checked.)</p><h3>Testing type conflicts across different mappings</h3><p>In the end, one of the most important testing dimensions was using different indices with various mappings in the same query. (Recall how, above, we had to deal with type conflicts to come up with a solid approach for PUNKs? It doesn’t end there; all kinds of type conflicts are more complex with <code>LOAD</code>.) Since we couldn’t automatically generate correct expected results, ES|QL’s test suite had to grow by adding more than 10,000 lines of CSV spec tests. Fortunately, adding such tests is a well-suited task for an AI agent, which has cut down the effort dramatically. (Of course, the test results were still reviewed by humans.)</p><p>All testing strategies together provided us with good confidence for the GA release of <code>unmapped_fields</code> with Elasticsearch 9.5.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-unmapped-fields-deep-dive</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-unmapped-fields-deep-dive</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Mappings]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Alexander Spies]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9f8079c053533ad8/6a8ee79f8658b748a0469342/image4.png" length="0" type="image/png"/>
    <pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[ES95: Adaptive Compression for Elasticsearch Time-Series Metrics]]></title>
    <description><![CDATA[ES95 is Elasticsearch 9.5's new adaptive time series codec that cuts @timestamp storage by 92% and floating point fields by up to 74%, with zero configuration.]]></description>
    <content:encoded><![CDATA[<p><em>The best compression strategy is the one that understands your data.</em></p><p>Observability workloads are storage-intensive by nature, and the composition of that storage determines both cost and query performance. <code>ES95</code> introduces adaptive compression: rather than applying the same encoding to every numeric field, it automatically selects the encoding that best matches each field's structure. The result is a 33.6% reduction in total doc-values storage, 19% to 74% reduction on floating-point gauge metrics and a 92% reduction in <code>@timestamp</code>. No configuration or migration required.</p><h3>Observability data is storage-intensive</h3><p>Data processing systems are rarely limited by how fast they can compute. They’re limited by how fast they can move bytes: off disk, across the network, and through the memory hierarchy. Compression is how a storage engine trades CPU time for memory bandwidth, spending comparatively cheap CPU cycles so fewer bytes have to travel through the parts of the system that are usually constrained. In a read-heavy system like Elasticsearch, that trade-off pays back every time data is queried, often long after it was written.</p><p>Storage size and query performance move together; fewer bytes on disk means fewer bytes to read on every range query, every aggregation and every dashboard load. Compression is not just about saving storage. Every byte that is never written is also a byte that never has to be read.</p><p>The right encoding depends on the structure of the values themselves, and the largest wins come from exploiting the structure already present in the data rather than squeezing an opaque stream of bytes. Few workloads expose that structure more clearly than observability metrics.</p><p>A single host reports hundreds of metrics every few seconds, including CPU utilization, memory ratios, request latencies, and network throughput. Multiply that by thousands of hosts across weeks of retention, and the bytes accumulate fast. Most of that volume is structured but not uniform: timestamps arrive at near-constant intervals from thousands of concurrent series, counters increase monotonically, while gauges like <code>23.47</code> or<code>1.15</code> are short decimal measurements.</p><p>A fixed compression approach cannot adapt to that variety. A timestamp column and a floating-point gauge column compress through fundamentally different techniques, but a codec that applies the same approach to both will necessarily handle one of them poorly. For most of Elasticsearch's time-series codec history, gauges were on the losing end of that trade-off.</p><h3>The structure the old codec was not built to exploit</h3><p>Elasticsearch stores time-series numeric values in <em>doc values</em>: a column-oriented structure where all values for the same field sit adjacent on disk. That adjacency makes compression possible: the codec compares consecutive values of the same field, finds patterns, and exploits them.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2386e2eb944a502a/6a85671fa8b3235ccccbf4f9/unnamed.png" alt="Row-oriented vs column-oriented storage in Elasticsearch time series indices showing how field values cluster on disk" /><p>The <a href="https://www.elastic.co/search-labs/blog/time-series-data-elasticsearch-storage-wins">time-series codec before ES95</a> applied the same fixed encoding to every numeric field: delta encoding followed by normalization, GCD (greatest common divisor) reduction, and bit-packing. Each encoding technique activated where it helped and skipped where it would not. For timestamps and integer counters, this approach was remarkably effective. For floating-point gauges, it could find almost nothing to work with.</p><p>The reason is how they’re stored. To support range queries, Elasticsearch stores floating-point values as integers that preserve numeric ordering. A change of 0.01 in a CPU percentage reading translates to a jump of trillions in that integer space. The codec sees those large jumps and has no strategy to further reduce their footprint. Storage stays near the original eight bytes per value.</p><p>The codec was doing the right thing with the representation it had, but the latter was chosen for querying, not compression, and the goals conflict at the bit level.</p><h3>The cost of a fixed format</h3><p>A compression stage for floating-point values was already on the roadmap, so the interesting part wasn’t the algorithm. The obstacle was architectural.</p><p>The previous codec baked its compression approach into the storage format. Adding a new encoding meant changing the meaning of existing bytes on disk, which forced a format migration, a rollout that can last weeks or months in large production clusters. Over time, that migration burden constrains codec development itself. The question stops being <em>Is this a good compression idea?</em> and becomes <em>Is it worth another format migration?</em> That rigidity limits the cadence of codec evolution and leads to missed compression improvements.</p><h3>The right encoding without configuration</h3><p><code>ES95</code> solves this at the architecture level for time-series indices. Each field's encoding is no longer baked into the format. It is selected automatically at write time, based on what the field mapping already declares: the field's name, its data type, and its metric role. Timestamps are encoded differently than counters. Counters are encoded differently than gauges. <code>ES95</code> encodes all of them, and it chooses the right strategy for each.</p><p>Users already tell Elasticsearch everything the codec needs to know. The mapping describes the data; the codec chooses the compression strategy.</p><p>Compression strategy is a codec concern, not a user concern.</p><p>The alternative would have been to expose per-field encoding selection as a configuration parameter, letting users opt in to better compression for specific fields. That would shift the burden of knowing which encoding fits which data type onto those least equipped to make that call and would guarantee that most deployments never see the benefit. <code>ES95</code> keeps that decision inside the codec, where it belongs. This matters most in managed and serverless deployments, where users expect the system to automatically make optimal storage decisions.</p><h3>The timestamp result nobody planned for</h3><p>With the adaptive architecture in place, the team set out to ship the planned float-compression algorithm. Before it arrived, the architecture proved itself by substantially improving compression for timestamps.</p><p>A time-series index is sorted first by its time-series identifier (<code>_tsid</code> constructed by the metric’s dimensions) and then by timestamp within each series. Timestamps on disk aren’t one smooth sequence; there are many smooth sequences laid end to end, one per series, with a large jump at every boundary where one series ends and the next begins.</p><p>The codec compresses data in fixed-size blocks without regard to those series boundaries. A block straddling a series boundary holds timestamps from two different series. The jump between them breaks monotonicity, reducing delta encoding effectiveness on blocks spanning different time series. A block that would otherwise compress to near-zero bits per value ended up needing nine or more, because bit-packing encodes every value in a block using the same fixed number of bits, so one large jump sets the cost for all of them.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd7576b4b4a9e1c2d/6a85677d1eb9e5448c2c3336/unnamed.png" alt="SplitDelta encoding splits compression blocks at series boundaries, cutting @timestamp storage by 92.4%" /><p>In the ideal case, every block belongs to a single series: timestamps increase at near-constant intervals, delta encoding captures the regularity, and bit-packing compresses the result to near-zero bits per value. With few series, boundary blocks are rare and the overhead barely registers. On an observability cluster ingesting millions of documents across thousands of series, that changes. The cost scales along two dimensions: series count and data density. More series means more boundary events; sparser series means multiple jumps packed into single blocks. In high-churn environments, both compound, and boundary blocks accumulate into a standing tax on the most-read field in any time-series workload. It’s why <code>@timestamp</code> storage grew faster than the data that produced it.</p><p>The fix was to detect series boundaries and treat each run as its own independent sequence. Instead of trying to encode across the jump between two series, which forces every value in the block to pay the storage cost of that one large jump, each run is compressed on its own terms. The boundary simply becomes a seam: the jump is never seen by the encoder on either side.</p><p>That encoding is called <code>SplitDelta</code>. <code>@timestamp</code> and monotonic long counters now use it by default. No format change. No migration. Existing segments retain legacy encoding.</p><p>On the high-cardinality <a href="https://github.com/elastic/rally">Elasticsearch Rally</a> benchmark, that single unplanned encoding cut counters storage by 20%–30% and <code>@timestamp</code> storage by 92.4%, from 1.03 GB to 79 MB. Gigabytes to megabytes, and no, that isn’t a typo.</p><p>The pluggable pipeline had already paid for itself. <code>SplitDelta</code>, which wasn’t part of the original plan, slotted in without a format change or migration before ALP even shipped.</p><h3>ALP: recovering the decimal that was always there</h3><p>Most floating-point metrics can be expressed as short decimals with no loss of accuracy: CPU utilization at <code>23.47</code>, load average at <code>1.15</code>. <a href="https://dl.acm.org/doi/10.1145/3626717">ALP</a> (Adaptive Lossless floating-Point compression) recovers that decimal structure from the floating-point representation, converting values into integers that the existing pipeline already handles well. <code>ES95</code> feeds ALP's output into the same mature integer compression pipeline used for timestamps and counters, extracting additional savings rather than treating ALP as a standalone encoding. Values that don’t fit ALP's model (such as irregular high-precision floats or special values) fall back to direct bit-packing or the original representation without degrading the rest of the block.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt056bdb51022428c9/6a85689f4acc96ce082471b1/unnamed.png" alt="ALP converts floating-point time series metrics to integers for compression through the Elasticsearch encoding pipeline" /><p><code>ALP</code> lets Elasticsearch treat floating-point metrics according to the structure they actually contain rather than the binary representation they happen to use. It’s applied automatically to double-valued gauge fields, selected by field type and metric role, through exactly the door the architecture had built for it.</p><h3>What the numbers say</h3><p>Here’s what the <a href="https://github.com/elastic/rally">Elasticsearch Rally</a> benchmark looks like on a high-cardinality workload containing 2.26 billion data points. Results are from an internal <code>tsdb-metricsgen</code> benchmark.</p><p><strong>Field or metric</strong></p><p><strong>Storage reduction (%)</strong></p><p><code>@timestamp</code></p><p><strong>−92.4%</strong></p><p><code>cpu.load_average.5m</code></p><p><strong>−74.3%</strong></p><p><code>system.cpu.utilization</code></p><p><strong>−63%</strong></p><p><code>memory.utilization</code></p><p><strong>-19%</strong></p><p>Total doc values</p><p><strong>−33.6%</strong></p><p>That overall 33.6% reduction deserves context.</p><p>A time-series index contains more than metrics. Every data point also carries the labels that identify the series: host names, IP addresses, regions, container IDs. Those dimension fields are stored as keywords. <code>ES95</code> doesn’t target dimensions.</p><p>On this benchmark, two dimension fields, <code>host.ip</code> and <code>host.mac</code>, accounted for 44% of doc-values storage after <code>ES95</code> ran. The 33.6% total reflects that mix. The per-field breakdown is the honest picture. Compression for dimension fields is an active area of work.</p><p>The per-field variation is the most convincing result. Some gauges shrank by nearly three quarters, while others moved by less than a fifth. That spread is direct evidence that <code>ES95</code> matches compression to the structure actually present in each field. A fixed encoding treats every field identically and misses most of those wins.</p><h3>Better compression without extra configuration</h3><p>The storage reductions from <code>SplitDelta</code> and <code>ALP</code> are the most visible results of <code>ES95</code>. The more consequential result is the architecture that produced them.</p><p>Before <code>ES95</code>, every new compression technique required a format evolution. That reality shaped which ideas were practical to pursue. Today, new encodings become implementation decisions inside the codec rather than migration projects. Existing data never needs to move, and users gain better compression on newly written data simply by upgrading Elasticsearch. <code>SplitDelta</code> and <code>ALP</code> are the first encodings to benefit from this architecture. They will not be the last.</p><p>Asking users to choose compression algorithms would only duplicate information Elasticsearch already has. There are no per-field compression parameters to tune, and no expert knowledge is required to get good storage efficiency. Different fields get different strategies because <code>ES95</code> understands what kind of data each field contains, not because a user configured it. As the codec evolves, those decisions evolve with it. The API does not.</p><p>In Elasticsearch Serverless, good defaults are part of the product. Users expect the system, not configuration, to make storage decisions. <code>ES95</code> is designed to honor that expectation: encoding that starts right and gets better over time.</p><h3>The compression was always there</h3><p><code>ES95</code> establishes a new standard for how time-series codec evolution works. New encodings become implementation decisions, not migration projects. Users get better compression on newly written data with every Elasticsearch upgrade.</p><p>The compression was already in the data. <code>ES95</code> just removed what was hiding it.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/time-series-database-compression-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/time-series-database-compression-elasticsearch</guid>
    <category><![CDATA[Index Data]]></category>
    <category><![CDATA[Operations]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Salvatore Campagna]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt27b1a5308344647f/6a8566caf9838a4c963bea55/unnamed.png" length="0" type="image/png"/>
    <pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[How Elasticsearch's batched query phase improves search performance at scale]]></title>
    <description><![CDATA[The batched query phase can cut search execution time in half by reducing transport overhead and better distributing reduction work across the cluster.]]></description>
    <content:encoded><![CDATA[<p>The batched query phase, which is live in Elasticsearch Serverless and released in Elasticsearch 9.5.0, changes how search is coordinated across the cluster. The coordinating node now batches all queries for each data node into one, instead of a separate transport request per shard. Data nodes can then do partial result reduction themselves, rather than shipping everything back to the coordinating node. In coordination-bound workloads, this can cut search execution time in half.</p><h2>How a search is executed in Elasticsearch</h2><p>Let’s start with some Elasticsearch basics. When a search request lands on an Elasticsearch node, that node is called the <em>coordinating node</em>for the search. By default, any Elasticsearch node can act as a coordinating node. The coordinating node determines which shards need to be searched based on the index or indices specified by the search. The nodes that those shards live on are called <em>data nodes</em>. A coordinating node can also be a data node, as it may host shards that are relevant to the search.</p><p>Elasticsearch performs the search in two primary phases: the <em>query phase</em> (also known as the <em>scatter phase</em>) and the <em>fetch phase</em>(also known as the <em>gather phase</em>). The query phase is responsible for going to the data nodes and executing the query on each shard. Each shard responds with a set of document IDs (just the IDs, no data) and an associated score for each. These results are reduced. Next, the fetch phase goes back out to the data nodes to fetch the document <code>_source</code> (the data). </p><p>What exactly is a <em>reduction</em>in Elasticsearch? A reduction turns per-shard results from the query phase into a single merged result for the client. Suppose a search asks for the top five hits in a three-shard index, according to some relevance score. Elasticsearch must then get the top five hits from each targeted shard. Why? Because it’s possible that one shard contains the global top five, or, more likely, that the top five docs are spread across shards. </p><p>If three shards are being searched, the coordinating node will have 15 <code>(docID, score)</code> pairs after the query phase. These results are reduced: The documents with the top five scores are kept and the rest thrown away. Then the fetch phase reaches back out to the data nodes to get the documents’ <code>_source</code> data, which Elasticsearch then responds with.</p><p>shard1_results = [(id: 231, score: 0.871), (id: 445, score: 0.812), (id: 88, score: 0.754), (id: 312, score: 0.701), (id: 567, score: 0.643)]</p><p>shard2_results = [(id: 847, score: 0.921), (id: 76,  score: 0.843), (id: 125, score: 0.783), (id: 438, score: 0.729), (id: 590, score: 0.668)]</p><p>shard3_results = [(id: 512, score: 0.887), (id: 289, score: 0.798), (id: 74,  score: 0.741), (id: 631, score: 0.682), (id: 405, score: 0.619)]</p><p>// Take the top five scores from above (that’s the “reduction”)
reduced_result = [(id: 847, score: 0.921), (id: 512, score: 0.887), (id: 231, score: 0.871), (id: 76, score: 0.843), (id: 445, score: 0.812)]</p><h2>What is the batched query phase?</h2><h3>Without batching (shard fan-out)</h3><p>To understand the batched query phase, we must first understand how the query phase worked without batching. The following diagram represents a three-node Elasticsearch cluster with 12 index shards. Suppose a client search request lands on Node 1. That makes Node 1 the coordinating node for the search.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd8b116819f43dc50/6a79aa014423e189ba9fa08f/image3.png" alt="" /><p>The coordinating node first <em>fans out</em> to all shards, performing the client’s query on each shard. Each shard query is facilitated by a <em>transport request</em>, which is a protocol used for communication between Elasticsearch nodes. Notice that in this diagram, each shard gets its own transport request. Node 1 holds shards itself, but no network request is needed to query those shards.</p><p>After querying the shards, the coordinating node must reduce them, as described in the previous section. At this point, the coordinating node is holding 12 shard results. Once it reduces them all, it can proceed with the fetch phase and then respond back to the client.</p><h3>With batching</h3><p>So what does the batched query phase change? The batched query phase first takes effect before dispatching the shard query transport requests. Now the coordinating node batches the shards it needs to query for each data node and requests them all at once.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt44a9f6586a927f2b/6a79aa2b2888390ea907d7c2/image2.png" alt="" /><p>Notice that only one transport request was made to Node 2 and one to Node 3. This is the first change in the batched query phase. Each transport request induces some overhead, so this change alone buys us our first performance improvement, reducing latency and CPU cycles spent on overhead.</p><p>The second change in the batched query phase has to do with reducing shard results. Now partial reductions occur on the data nodes themselves. For instance, after querying shards 2, 5, 8, and 11, Node 2 then reduces those results. The coordinating node, upon receiving partially reduced results from Node 2 and Node 3, must then perform afinal reduction. This is the second major enhancement we get from the batched query phase: the spreading out of reduction work across the data nodes. This reduces memory pressure on the coordinating node, which we’ll see measured later in the benchmarking section.</p><h2>The benefits of batching</h2><p><em>Fan-out</em> (no batching) is how search worked for a very long time. It has advantages: each shard is a separate request that returns and can be retried independently from all the others. With many shards involved, though, each shard request causes overhead due to many round trips going between the coordinating node and the data nodes.</p><p>Also, the coordinating node doesn’t have enough information to be able to determine the pace at which to send requests to each data node. It sends a maximum of five concurrent requests per data node by default, where five is a bit of a magic number which allows for some parallelism, while at the same time preventing a single query from taking over an entire data node. At the same time, data nodes also don't make distinctions between the different shard requests they receive, for instance based on what parent search request they belong to. In reality, if all shard requests involving a search request are presented in one batch to each data node, the data node can then look at its internal state and adapt its pace dynamically, removing the five concurrent shard requests artifical limit mentioned above.</p><p>The batched query phase results in several benefits, including:</p><ul><li><p>Increased efficiency, thanks to fewer round trips: less CPU spent on transport overhead and fewer bytes going through the transport layer.</p></li><li><p>Spreading out the load of reductions: the coordinating node was previously the bottleneck for reductions, and now data nodes share the work.</p></li><li><p>Better resource usage: we may be able to better max out the data node’s CPUs.</p></li></ul><p>The average search against many shards can now be served much quicker, with lower latency and higher throughput.</p><h2>Batched query phase benchmarks: Latency and memory usage</h2><p>To measure the benefits of the batched query phase, we ran a couple of benchmarks. The first displays the benefit of reduced transport overhead. This benchmark was built off the “many-shards-quantitative” <a href="https://elasticsearch-benchmarks.elastic.co/">nightly benchmark</a>. It runs on a three-node Elasticsearch cluster, running batches of searches targeting 1,000, 5,000, and 20,000 shards. The queries are <code>match_all</code> queries with <code>size: 0</code>. That means querying the shards themselves is effectively a no-op. This benchmark is meant to isolate the work of coordinating the search across the cluster.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd4a4a7aadc1632a8/6a79aa47dcb437094d2cf8da/image1.png" alt="" /><p>We ran the same benchmark with and without the batched query phase (by toggling the cluster setting <code>search.batched_query_phase</code>). At 5,000 shards, the batched query phase makes searches twice as fast. By 20,000 shards, it’s 2.2x faster. These results are a ceiling for what a realistic workload can expect to benefit; any real query work at the shard level will dilute the overall result. However, the gains are real, and the <em>coordination work</em> of your queries will benefit as shown. Note that this first benchmark did not attempt to measure any gain brought by spreading reductions across data nodes, as opposed to performing them only on the coordinating node.</p><p>Next, we benchmarked the benefits of the batched query phase on large reductions. In our benchmark, we ran a large terms aggregation over a data set called <code>http_logs</code> (which can be found in our <a href="https://github.com/elastic/rally-tracks">rally-tracks</a> repo). This data set has 247 million documents, which we indexed in seven indices each with 100 shards, again in a three-node Elasticsearch cluster. We ran a single <a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-bucket-terms-aggregation">terms aggregation</a> query for the <code>clientip</code> field with <code>size: 1000</code> and <code>shard_size: 50000</code>. That means we’re asking for the top 1,000 terms, although each shard will return 50,000 buckets to be reduced to that 1,000. Requesting so many terms from each shard increases precision, but has a cost in terms of memory usage and reduction overhead, which helps highlight the gain provided by spreading incremental reductions across data nodes.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt978ed04cf2406b0e/6a79aa5a8bcd807043261d4b/image4.png" alt="" /><p>The <em>memory used</em> is measured by something called a <em>circuit breaker</em>, which accounts for memory usage in Elasticsearch. This field can be found in the <code>/_nodes/stats</code> response under <code>nodes.&lt;node_id&gt;.breakers.request.estimated_size_in_bytes</code>. It estimates the memory in use for processing in-flight search requests.</p><p>The memory benchmark shows how the work of reductions is now spread across the cluster. In blue is the memory used with <code>search.batched_query_phase: false</code>. We can see that the node elasticsearch-0 is the coordinating node, as it bears the entire reduction load itself. This effect is no longer the case with <code>search.batched_query_phase: true</code> in red. The node elasticsearch-1 is the coordinating node, but it uses far less memory. That’s because the data nodes elasticsearch-0 and elasticsearch-2 do reductions themselves, reducing the results before they’re returned to the coordinating node.</p><h2>Summary</h2><p>Batching shard queries per data node reduces transport overhead and allows each data node to perform partial results reduction locally. The result is lower latency for searches targeting many shards and better distribution of work across the cluster for reduction-heavy workloads. At the same time, the new model works well for all scenarios and opens the door for potential future enhancements which we’re excited to pursue, such as:</p><ul><li><p>Optimize out-of-the-box resource usage by introspecting data nodes' activity and adjusting the pace accordingly, which would replace the five concurrent shard requests per data node artifical limit.</p></li><li><p>Assign priorities to search requests and let each data node process shard requests accordingly.</p></li><li><p>Reduce the number of roundtrips further by folding the `can_match` phase into the query phase. Can match is a separate "batched" roundtrip used to shortcut the query against shards that can't possibly match based on index statistics. Now that the query phase is batched, both rounds can be executed in one go.</p></li></ul><ul><li><p>Discontinue support for minimize roundtrips in cross-cluster searches in favour of batched query execution; in hindsight, minimize roundtrips achieves a similar goal but is specific to cross-cluster execution, while batched execution is applicable to every search. </p></li></ul>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-batched-query-phase</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-batched-query-phase</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Vector Database]]></category>
    <dc:creator><![CDATA[Ben Chaplin,Luca Cavanna]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltec0235a331fb0c48/6a79a9e30da673867657b299/image5.png" length="0" type="image/png"/>
    <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Why Elasticsearch is becoming a columnar database]]></title>
    <description><![CDATA[Elasticsearch is becoming a first-class columnar database. Columnar Mode ships in 9.5, storing data once alongside the existing modes and cutting storage footprints while speeding up analytical queries.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch is becoming a first-class columnar database. In 9.5, we’re introducing <strong>Columnar Mode</strong>, a new index mode that stores data once, in columnar form, with no redundant copies and no indexes the workload doesn't need.</p><p>Columnar Mode targets the workloads where data is written in volume, queried analytically, and retained for a long time: logs and observability, security telemetry, metrics and traces, business analytics on operational data, and AI retrieval at scale. For these workloads, it means <strong>meaningfully smaller storage footprints</strong> on day one, and the foundation for faster ingest, faster analytical queries, and longer retention as the surrounding work matures.</p><p>Columnar Mode ships <strong>alongside</strong> the existing modes, not in place of them. There are no changes to APIs, dashboards, applications, or downstream integrations. It's a new index mode that you can adopt where it helps, on the data where it fits.</p><p>The result is that Elasticsearch becomes a <strong>world-class search engine and columnar analytics engine</strong>, on the same data, in the same cluster, under the same operational umbrella. This is how Elastic stays the most useful place to put operational data as the economics of that data are rewritten by columnar architectures.</p><p>The rest of this post explains why we're doing it, along with what it changes and what it doesn't.</p><h2>Why Elasticsearch is adding a columnar mode</h2><p>Almost every general-purpose data system worth caring about has, over the last 15 years, made the same architectural decision: When the job is to read, aggregate, and reason over large volumes of data, you organize that data <strong>by column</strong>, not by row or document.</p><p>Elasticsearch is now making that decision, too</p><p>In addition to its document model, search heritage, and place at the heart of observability, security, and search applications across the industry, it’s now adding a second way to think about data inside the same platform. We're calling it <strong>Columnar Mode</strong>. It positions Elasticsearch to be the only major platform that does <strong>search</strong>, <strong>retrieval</strong>, and <strong>analytics</strong> at the same level, on the same data, with one operational story.</p><p>To get there we need to point out three things: One, how Elasticsearch came to be what it is today and why the document model was the right call. Two, how the rest of the data world, in parallel, converged on a completely different shape and why that movement has been so successful. And three, why these two worlds are now converging inside our users' systems and why putting columnar inside Elasticsearch is the right answer.</p><h2>How Elasticsearch became a document database (and why that was the right call)</h2><p>To understand where we're going, you have to understand where we started.</p><p>Sixteen years ago, when Elastic’s founder Shay Banon wrote the first version of Elasticsearch, the data world was in the middle of a quiet revolution. JSON was eating the wire. Web applications were sending and receiving data not as carefully typed rows in clearly defined tables, but as flexible, semi-structured documents (a customer order, a chat message, a product listing, a log line, a recipe), each a self-contained little world of fields, sub-fields, arrays, and nested objects.</p><p>The relational databases of the time were built around a very different assumption. Define a strict schema up front. Decompose your data into normalized tables. Use joins at query time to reassemble it. This worked beautifully for the use cases it was designed for (transactional systems, banking, enterprise resource planning systems), but it created enormous friction for the new generation of web and mobile applications. Adding a field meant a migration, and storing variable-shaped data meant either a brittle schema or a tangle of nullable columns.</p><p>In response, a new category of system emerged: document databases. The bet was simple. Take the data as it arrives, and store it in something close to its native shape. Don't make developers fight the database to express their domain. MongoDB became the flag-bearer for this idea on the operational side, and Elasticsearch carried the same banner on the search side.</p><p>For Elasticsearch, the document model was not just a developer-experience choice. It was a deep architectural commitment. The engine was designed around the idea that a record is a document, a document has fields, fields have types, and you should be able to put almost anything in and get useful behavior out (full-text search, structured filtering, sorting, aggregations, ranking) without having to model your data perfectly in advance.</p><p>This is what we mean when we talk about Elasticsearch's "document behavior":</p><ul><li><p>The engine remembers what you sent it. The original record is stored and can always be returned exactly as it was.</p></li><li><p>Every field is, by default, made queryable in multiple ways: optimized for fast exact lookups, fast text search, fast range queries, fast sorting, and fast grouping.</p></li><li><p>Schemas are flexible. New fields are absorbed at write time. The cost of getting your data model "wrong" is low.</p></li><li><p>The unit of work is the document. Indexing, retrieval, and most APIs are framed in terms of "given a document, do X" or "given a query, return documents."</p></li></ul><p>This is a model born from search. Apache Lucene, the storage library Elasticsearch is built on, was designed to make one workload remarkably fast: Take a query, find the small number of documents that match it, rank them by relevance, and return them. To do that well, you build <strong>inverted indexes</strong>: structures that turn the question around, so instead of asking "What's in this document?" you can ask "Which documents contain this term?" and get an answer in milliseconds. You also keep the original document around, because once you've found it, you usually want to show it or retrieve more information from it.</p><p>For the workloads Elasticsearch grew up serving (site search, log search, application search, security event lookup, "find this in a haystack") this was, and still is, the right shape. There's a reason Elasticsearch is used in production at hundreds of thousands of organizations worldwide.</p><p>But the document model has a built-in cost, and that cost is the starting point for everything that follows.</p><p>To keep its promise of fast search, flexible schemas, and faithful storage of the original record, the engine ends up storing data multiple times, in multiple shapes, each optimized for a different question. The original document is stored, and the text in each field is indexed for search. The values in each field are also stored separately, in a per-field columnar store known as <em>doc values</em>, so that they can be aggregated, sorted, and grouped. The system maintains all of these in parallel, by default, on every field, because at write time it doesn't know which of these capabilities you'll need at read time.</p><p>For the kinds of datasets Elasticsearch was originally designed for (relatively rich, relatively low-volume documents where every read is valuable), this trade-off is excellent. You pay a modest storage and ingestion tax for an enormous capability surface.</p><p>For the kinds of datasets Elasticsearch increasingly stores today (billions of log lines, trillions of metric points, oceans of telemetry, most of it written once and read rarely), that trade-off starts to look very different.</p><p>That’s the friction we’re setting out to resolve.</p><h2>The other story: How the data world went columnar</h2><p>While Elasticsearch was perfecting the document model, a parallel revolution was happening in analytics. It started in academia, became commercial, and is now the default architecture for almost every system built to read and reason over large datasets. To position Columnar Mode clearly, we have to tell that story.</p><p>The bet behind columnar storage is so simple it almost sounds like a trick. Instead of storing your data <em>row by row</em> (record one, record two, record three…), store it <em>column by column</em> (all values of <code>timestamp</code>, then all values of <code>host_name</code>, then all values of <code>status_code</code>…).</p><p>That single decision changes everything.</p><p>It started in the early 1990s, with research systems like MonetDB out of the Centrum Wiskunde &amp; Informatica in Amsterdam. It was sharpened in the mid-2000s with C-Store, a project led by Michael Stonebraker at MIT, which became the basis for Vertica. Sybase IQ carried the same idea into the early commercial market. By the early 2010s, every serious analytics vendor on earth had a columnar story. Today, the lineage runs through Amazon Redshift, Google BigQuery<strong>,</strong> and Snowflake in the cloud warehouse world; through open file formats, like Apache Parquet and Apache ORC that have become the lingua franca of the data lake; through in-memory standards like Apache Arrow; and through fast operational columnar engines, like ClickHouse and DuckDB.</p><p>Why has this approach won so completely?</p><p>Because when your job is to read and reason about a lot of data, columnar storage turns nearly every dimension of cost and performance in your favor.</p><p><strong>You only read what you need.</strong> Real analytical queries almost never want every field; they want three fields out of 50, or aggregate one field out of 200. A row-shaped store has to walk past all the other fields on the way to the ones you care about. A columnar store reads only the columns the query touched. On wide datasets, this routinely cuts the I/O of a query by 90% or more.</p><p><strong>You compress dramatically better.</strong> When you put all the values of one field next to each other, those values are, by definition, of the same type, often with repetition and similar structure; for example, a column of timestamps from the same hour, a column of status codes that are almost always 200, or a column of hostnames drawn from a small set. Compression algorithms thrive on this kind of homogeneity. Where row-shaped stores typically achieve 1.5–3x compression, columnar systems routinely achieve 5–10x, while on low-cardinality fields, it's common to see 20–30x. That’s not just a tuning improvement; that’s a different economic regime.</p><p><strong>You can skip vast amounts of data without reading it.</strong> Because data is organized in blocks of values from the same column, the engine can carry tiny pieces of metadata for each block (minimum, maximum, count, distinct values) and use them to prune entire blocks at query time. Looking for errors in the last hour? Skip every block whose timestamp range is older. Looking for a specific hostname? Skip every block where it doesn't appear. The system avoids work it doesn't need to do by knowing in advance that the work is pointless.</p><p><strong>You can fully leverage the power of modern CPUs.</strong> Modern processors are designed to operate on long, regular, predictable arrays of values; that's where the cache, pipeline, and vector units do their best work. A column is exactly such an array. Columnar engines run queries by passing batches of values through tight, vectorized operators, instead of looking up one record at a time. The speed gains aren’t incremental: Across academic benchmarks reaching back to Abadi et al.'s seminal <em>Column-Stores vs. Row-Stores</em> (SIGMOD 2008) and the comprehensive surveys that followed, columnar engines have consistently shown one to three orders of magnitude advantage over row-shaped systems on analytical workloads.</p><p><strong>You delay materializing the original record for as long as possible.</strong> Because the data is already organized by what queries actually do (read this column, filter on that column, aggregate this column), there's no need to reconstruct a full row until you absolutely have to. Most of the time, you don't have to at all, especially if you run analytical queries that only care about aggregating a few columns.</p><p>The cumulative effect of these properties is the reason every cloud data warehouse is columnar and why every modern data lake stores its files in Parquet or ORC. It’s also the reason new operational analytics engines have, almost without exception, chosen this shape. It’s the structural answer to the question "How do you make queries over large data cheap and fast?" and the answer turns out to be the same almost everywhere.</p><p>We should also note what columnar systems give up. They aren’t designed for "Give me one specific record and update it in place" or for transactional consistency at the level of individual rows. In general, they aren’t designed to do full-text relevance ranking the way a search engine does. They’re optimized for a different shape of workload, and that's the point. Different workloads, different defaults.</p><p>For a long time, the world worked because users could pick one tool for one workload and another tool for the other. That world is ending.</p><h2>Why search and analytics are now converging</h2><p>Three things have happened in the last five years that change the calculus.</p><p><strong>First, the volume of data users generate has exploded.</strong> Driven by microservices, containers, OpenTelemetry, AI workloads, and security telemetry, industry estimates put observability and log data ingest at large enterprises in the multiple-terabytes-per-day range, with annual growth rates that have run well into triple digits for several years.</p><p><strong>Second, the cost of storing and querying that data has become a top-line concern.</strong> Industry surveys and practitioner reports consistently flag observability and telemetry as one of the largest and fastest-growing line items in modern infrastructure budgets, and a meaningful share of that spend goes toward data that’s written, retained, and then never queried.</p><p>These first two pressures have already produced their answer in the market: Modern columnar systems built for analytics (most visibly ClickHouse on web logs and the cloud warehouses on broader analytics) have redefined what users expect to pay per terabyte stored and per query run. Elasticsearch is now meeting that bar without giving up what makes it Elasticsearch.</p><p><strong>Third, the boundary between search workloads and analytical workloads has dissolved.</strong> Virtually every modern operational use case wants both: find me this specific event, and then aggregate everything around it; show me this log line, and then chart the trend that produced it; retrieve this document, and then group what else matches.</p><p>The users who put their logs, metrics, traces, security events, and search corpus into Elasticsearch aren’t asking us to be only a search engine or only an analytics engine. They’re asking us to be one engine that does both, on the same data, with one operational story, at the cost basis they expect from a modern columnar system.</p><p>That’s a high bar, and meeting it requires us to change something fundamental about how Elasticsearch organizes data.</p><p>For a long time, Elasticsearch has had columnar storage <em>underneath</em>. We’ve stored per-column data in our engine since 2013, but we’ve always treated it as a secondary structure layered on top of a document model. That made sense when the document model was the source of truth and the columnar layer was an optimization. It makes less sense when, for an enormous and growing share of our users' data, the columnar shape <em>is</em> the source of truth and the document model is the optimization most don’t need.</p><p>This is what Columnar Mode is for.</p><h2>Columnar Mode: Elasticsearch with a second way of thinking about data</h2><p>The decision we’re making is deliberately conservative in its surface area and ambitious in its impact.</p><p>We aren’t building a new product, asking users to migrate, or changing the API, the query language, the management surface, the integrations, the visualization layer, or anything else that touches their operational reality. Elasticsearch is still Elasticsearch.</p><p>What we are doing is introducing a new mode that users can apply, index by index, that turns Elasticsearch into a first-class columnar system for the data where that's the right shape.</p><p>In Columnar Mode, the engine flips its defaults:</p><ul><li><p><strong>Data is stored once, in our columnar store (doc values).</strong> No parallel copy of the original document and no secondary structures built by default. Each field is responsible for its own storage, and the engine doesn't pay for capabilities the workload doesn't need.</p></li><li><p><strong>Search indexes are built only where they earn their keep.</strong> The message field of a log is still indexed by default to enable free-text search. A numeric field used only in aggregations means no index, extra storage, or additional write-time cost. The engine becomes leaner by default and lets you opt in to capability where it's worth paying for.</p></li><li><p><strong>The original record can be regenerated on demand</strong> from the column store, instead of being kept as a redundant copy. Users who want to keep the stored copy for query convenience can; and users operating at petabyte scale can opt to drop it entirely for further storage savings.</p></li><li><p><strong>The data model is genuinely columnar.</strong> Fields are flat key/value pairs, not nested object trees. Multi-valued fields, cardinality, and nullability are first-class mapping concepts, the same primitives that make pure columnar systems efficient.</p></li><li><p><strong>Specialized profiles ship on top.</strong> The first one is <strong>Columnar Logs</strong>, a logs-oriented variant of the mode with indexing enabled on log messages and the right defaults for time-ordered data. We’ll follow with profiles for other workloads, including, eventually, a columnar profile for vector retrieval, using the same building blocks.</p></li></ul><p>All of this is new for Elasticsearch though not conceptually new in the industry.</p><p>The reason this matters is that we’re doing this <strong>without leaving anything behind</strong>. The same query that runs against a Columnar Mode index runs, unchanged, against a document-mode index. Plus, the same dashboards, agents and integrations, Service Level Objectives (SLOs), alerts, rules, and machine learning jobs work as usual. Users who want the document behavior can always keep it on the indexes where it makes sense. Users who want columnar efficiency on the indexes where <em>that</em> makes sense (typically, the largest indexes in their cluster) get it.</p><p>This is the difference between us and a pure columnar engine. A pure columnar engine does one of these jobs and asks you to bring another system for the other. Elasticsearch is one platform that does both.</p><h2>What Columnar Mode is for</h2><p>Columnar Mode is built for workloads where data is written in volume, queried analytically, and retained for a long time, rather than a universal default. These work categories include:</p><p><strong>Logs and observability at scale.</strong> Users running tens or hundreds of terabytes of logs per day spend most of their observability budget on storage and on the aggregation queries that power dashboards, SLOs, and rate calculations. Columnar Logs is the first specialized profile precisely because this is where the impact lands hardest. At petabyte scale, Columnar Mode changes what is economically possible, that is, longer retention, more raw fidelity, and fewer compromises forced by cost.</p><p><strong>Security event stores and threat hunting.</strong> Security telemetry has the same shape as logs, but with a heavier emphasis on faceted exploration, ad hoc correlation, and deep historical lookups for indicators of compromise. Columnar Mode keeps full-fidelity storage affordable while preserving the search-relevance behavior security analysts need for lookup and pivoting, using the same engine to handle both jobs.</p><p><strong>Metrics, traces, and the unified observability substrate.</strong> Time-series data is already columnar in TSDB, Elasticsearch’s time series database. A future columnar profile will share the same storage building blocks across logs, metrics, and traces, letting all three sit in one engine on the same substrate without any of them compromising on workload-specific behavior. The unified observability platform becomes economically practical, not just architecturally elegant.</p><p><strong>Operational and business analytics on application data.</strong> Beyond observability, this means dashboards over millions of orders, transactions, user events, and Internet of Things (IoT) readings. Many of the workloads that have historically pushed users to ship the same data into a separate analytical warehouse, purely to make the queries affordable, can, with Columnar Mode, stay where the data already is.</p><p><strong>AI retrieval at scale.</strong> Vector workloads are inherently columnar. A future columnar profile for vector retrieval will combine the dense storage efficiency of the general mode with the indexing structures retrieval needs, putting retrieval augmented generation (RAG) and semantic-search workloads on the same cost basis as every other shape of data in the cluster.</p><p>The thread running through all of these is that Columnar Mode is the right answer when data is append-mostly, queries are analytical, and access patterns are per-column. For the workloads where those conditions don't hold, like search-first applications where the document is the unit of value, transactional flows that update individual records, or anything that depends on rich document structure as user-facing semantics, the document-oriented modes remain the right default, and they’ll continue to be invested in.</p><p>From the start, the principle has been that different workloads require different defaults, and Columnar Mode is what makes that real.</p><h2>What changes for Elasticsearch users with Columnar Mode?</h2><p>For Elasticsearch users, here’s what changes:</p><ul><li><p><strong>Costs go down meaningfully.</strong> Columnar Mode only builds inverted indexes where they earn their keep and stops storing data multiple times on every field. The result is a smaller storage footprint for the same workloads. Users can choose between two shapes of the mode: Columnar Logs keeps an inverted index on the message field, where the bulk of log queries actually look; the pure Columnar Mode variant drops that default entirely on string-type fields. A follow-up technical deep dive will publish the benchmarks and exact storage savings. The savings compound at scale and show up in operational reality immediately on self-managed clusters in the disks and machines the cluster no longer needs, and on Elastic Cloud in the bill.</p></li><li><p><strong>The architecture pays off on reads and writes, too.</strong> The same choices that make Columnar Mode storage-efficient (data stored once, indexes only where they earn their keep, and vectorized execution over columns) are the ones that shape its behavior on the read and write path, as well. The workloads that benefit most are the ones that dominate observability, security, business intelligence, and AI retrieval.</p></li><li><p><strong>The model gets simpler.</strong> Columnar Mode is opinionated by default. The right behavior for analytical workloads is what you get out of the box, like what we’ve done for metrics with TSDB, but for all types of data. The configuration surface contracts in places it should have contracted years ago.</p></li><li><p><strong>The path is nondisruptive.</strong> Users can adopt Columnar Mode index by index. The cluster, applications built on top of Elasticsearch, and dashboards don't change. The model is additive: a new way of behaving that’s available where it helps, alongside the behavior they already trust.</p></li></ul><p>The result is an Elasticsearch that’s leaner and faster, with no changes where it has been working all along.</p><h2>What doesn't change when you turn on Columnar Mode</h2><ul><li><p><strong>The document-oriented modes aren’t deprecated.</strong> Every existing mode remains available, supported, and invested in. Users who depend on document behavior (typical search applications, application search, security workflows, and anything where the original record's structure is the point) keep that behavior.</p></li><li><p><strong>Existing modes get faster, not slower.</strong> The work behind Columnar Mode is, at its core, work on doc values, Elasticsearch's columnar store, which every mode uses under the hood. Compression, encoding, query-execution, and vectorization improvements that ship with Columnar Mode benefit document-mode indexes alongside columnar ones for fields that reside on doc values. LogsDB, TSDB, and standard indexes get meaningfully faster analytical queries as a side effect of this work, with no user action or migration required. Investing in Columnar Mode is investing in every mode.</p></li><li><p><strong>The APIs don’t change.</strong> Every interface a user or partner integrates with (the REST APIs, query languages, management UI, data collection agents, integrations, and SDKs) keeps working exactly as it does today. Columnar Mode is an index setting, not a fork in the product, and downstream components don’t need to learn anything new.</p></li><li><p><strong>Search relevance, vector search, and semantic retrieval aren’t affected.</strong> These remain first-class capabilities of the engine, and they benefit from many of the same query and indexing performance investments that Columnar Mode rides on.</p></li><li><p><strong>There’s no migration cliff.</strong> Existing indexes keep working, and new indexes can be created in whichever mode fits their workload. Users move at the pace that makes sense for their organization.</p></li><li><p><strong>Elastic Cloud Serverless gets columnar the same way.</strong> On Serverless, where users manage projects rather than indices, Columnar Mode becomes part of the project type configuration. Observability and log-heavy project types will move to columnar defaults as the mode matures, with no user-facing change in how projects are created or managed.</p></li></ul><p>This matters because the temptation, when a company makes a big architectural bet, is to oversell it as the new gospel and undersell the engine's existing strengths. We don't have to make that trade. Elasticsearch's existing modes aren't legacy; they’re a mature, production-hardened document-oriented search engine that has been refined by more than a decade of use at scale. Columnar Mode is what they've been missing for the workloads they were never designed for.</p><h2>Why columnar storage matters for the future of Elasticsearch</h2><p>For most of the last decade, the dominant pattern has been fragmentation; for example, a search engine for search, a warehouse for analytics, a time-series database for metrics. This also includes a log aggregator for logs, a vector database for AI retrieval, and a graph database for relationships. Each one has its own ingestion pipeline, query language, operational story, and bill. The complexity tax this imposes on users is enormous, and the spend it generates for the industry is even larger.</p><p>The pattern that’s starting to replace it is <em>convergence</em>, the realization that the right architecture for a modern data engine is one that takes data once, stores it efficiently, and exposes it to whatever workload the application happens to need, whether that’s search, retrieval, aggregation, ranking, or vector similarity. That convergence is the most important trend in the data infrastructure industry, and it’s the trend on which Elasticsearch's future depends.</p><p>Columnar Mode is our central move in that direction, along with the work happening alongside it on indexing throughput, Elasticsearch Query Language (ES|QL, Elasticsearch’s analytical query language), vector retrieval, and observability and security as integrated solutions. It’s the architectural decision that unlocks the others, because without a first-class columnar layer, Elasticsearch's economics simply don’t compete on the workloads where convergence is most valuable.</p><p>With Columnar Mode, Elasticsearch becomes the only major platform in the industry that can credibly be <strong>the best search engine you can run</strong> and <strong>a competitive columnar analytics engine</strong> on the same data, in the same cluster, under the same operational umbrella. Although there will always be a specialized system that wins a specialized benchmark for individual workloads, Elasticsearch is where operational data belongs when it spans more than one shape and workload.</p><p>That’s the bet, and we’re confident it’s the right one.</p><h2>The bottom line on Elasticsearch as a columnar database</h2><p>Sixteen years ago, the right call for a search engine was to behave like a document store. We made that call and built one of the most widely used data platforms in the industry.</p><p>Today, the right call for the next decade of operational data (observability, security, AI retrieval, application search, business analytics) is to give that same engine a columnar way of thinking as a first-class mode users can choose, on the data where it fits, without giving up any of the things they already rely on us for.</p><p>That’s what Columnar Mode is. It’s the most important architectural change we’ll make to Elasticsearch this decade, and it’s the move that makes Elasticsearch the engine our users will still be reaching for when their current tool of choice has been displaced.</p><p>We’re going columnar. The reasoning is sound, and the path is nondisruptive. Columnar Mode reaches technical preview in Elasticsearch 9.5 and general availability (GA) in Elasticsearch 9.6, and every product team at Elastic is building on top of this foundation. We aren’t waiting for the future of data systems to arrive; we’re building it.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-columnar-storage</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-columnar-storage</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Analytics]]></category>
    <category><![CDATA[Columnar]]></category>
    <dc:creator><![CDATA[Yannis Roussos]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt52c0c0ae90e62a4d/6a55eb9f28aa67a917e60bba/0efd01fd06a0b70a30b1ec74c4995c4acadf98e7-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Your compliance posture just got an upgrade: Elasticsearch now supports FIPS 140-3]]></title>
    <description><![CDATA[Elastic 9.4 brings FIPS 140-3 support for Elasticsearch and Kibana to GA. Here's what changes for federal, defense and regulated deployments, and how to migrate from 140-2.]]></description>
    <content:encoded><![CDATA[<p><strong>The latest National Institute of Standards and Technology (NIST) cryptographic standard is fully supported in Elastic 9.4, so your compliance posture and your software can move forward together.</strong></p><p><a href="https://csrc.nist.gov/pubs/fips/140-3/final">FIPS 140-3</a> support for Elasticsearch and Kibana is generally available in Elastic 9.4 for self-managed deployments. NIST has stopped accepting new FIPS 140-2 submissions, with existing certificates winding down in September 2026. Every layer of your infrastructure needs to catch up, and your search and analytics platform is no exception. Federal programs, defense integrators and regulated enterprises are actively moving procurement requirements to 140-3. With Elastic 9.4, your stack can answer "yes."</p><h2>Two problems, one deadline: why FIPS 140-2 is no longer enough</h2><p>If you've been running Elasticsearch in FIPS 140-2 mode, you already know the value of having a compliant Elastic Stack. But two pressures are converging:</p><ul><li><strong>The standard is sunsetting.</strong> FIPS 140-2 certificates are being phased out. Procurement officers, auditors, and authorization bodies are shifting their requirements to 140-3. A stack that only supports the old standard is a stack with an expiration date on its compliance story.</li><li><strong>Auditors don't accept "close enough."</strong> It's not sufficient to run FIPS-approved algorithms. Your application layer needs to be explicitly configured for FIPS mode, with nonapproved algorithms disabled and cryptographic boundaries clearly documented. Partial compliance is noncompliance.</li></ul><p><em>Can't I just put Elasticsearch behind a FIPS-compliant load balancer or run it on a FIPS-hardened OS?</em> No, and here's exactly why: FedRAMP's network security requirements apply to cryptographic operations at the application layer, not just the network boundary. Intranode communication, keystore encryption, and credential hashing all need to happen inside a validated module. A compliant perimeter around a noncompliant application layer isn't a FIPS deployment; it's an audit finding waiting to happen.</p><h2>What FIPS 140-3 mode does in Elasticsearch 9.4</h2><p>When you flip FIPS mode on, Elasticsearch and Kibana restrict every cryptographic operation to FIPS-approved algorithms and delegate all of it to the validated module in your runtime. Here's what that covers and how.</p><ul><li><strong>High-grade security Transport Layer Security (TLS) everywhere, no exceptions.</strong> Node-to-node transport, the REST API over HTTPS, Kibana talking to Elasticsearch: All of it uses only FIPS-approved cipher suites. Noncompliant suites aren't deprioritized. They're rejected.</li><li><strong>PBKDF2 replaces bcrypt for password hashing.</strong> Bcrypt isn't FIPS-approved, so in FIPS mode, Elasticsearch switches to PBKDF2 for the native realm, file realm, and any stored credentials. Your users and service accounts stay protected with an algorithm that the auditor won't flag.</li><li><strong>Keystore encryption stays inside the boundary.</strong> Secrets in the Elasticsearch keystore, API keys, repository credentials, and encryption keys for Kibana's saved objects are wrapped with FIPS-approved key derivation and encryption. No gaps between <em>the cluster is compliant</em> and <em>the secrets are compliant</em>.</li><li><strong>You supply the cryptographic module.</strong> FIPS 140-3 support in Elasticsearch is built on the <a href="https://www.bouncycastle.org/fips-java/">Bouncy Castle FIPS Java API 2.0</a>, a FIPS 140-3 validated cryptographic module that runs inside the Java Virtual Machine (JVM). Elasticsearch delegates all cryptographic operations to Bouncy Castle; it doesn't implement its own crypto functions. Elasticsearch uses your own FIPS 140-3 JVM and Bouncy Castle FIPS provider, giving you full control over your cryptographic boundary and module versioning.</li></ul><p>Kibana takes a different path, running in a Node.js environment configured with a FIPS-compliant OpenSSL 3 provider. Together, both components operate within clearly defined cryptographic boundaries clean enough to diagram for an auditor.</p><h2>Who needs FIPS 140-3 support in Elasticsearch</h2><ul><li><strong>Federal and defense teams</strong> operating Elastic inside FedRAMP boundaries, Cybersecurity Maturity Model Certification–scoped (CMMC-scoped) environments, or Defense Information Systems Agency (DISA) Security Technical Implementation Guide–hardened (STIG-hardened) infrastructure. You can now upgrade to 9.x without punching a hole in your authorization documentation. Your Authorization to Operate (ATO) package references FIPS 140-3, not a soon-to-expire 140-2 certificate.</li><li><strong>Financial services and healthcare organizations</strong> where Payment Card Industry Data Security Standard (PCI DSS), Sarbanes‑Oxley Act (SOX), or Health Insurance Portability and Accountability Act (HIPAA) audits ask how your search infrastructure handles cryptography. FIPS mode gives your compliance team a one-word answer instead of a three-paragraph explanation.</li><li><strong>Anyone fielding </strong><em>Do you support FIPS 140-3?</em><strong> in a procurement questionnaire.</strong> That question is showing up in enterprise requests for proposal (RFPs), partner security assessments, and insurance underwriting checklists. With 9.4, the answer is <em>yes</em>.</li></ul><h2>Migrating from FIPS 140-2 to FIPS 140-3</h2><p>If you're running FIPS 140-2 on Elastic 8.x today, you don't have to jump to 9.x on day one. <strong>Elastic 8.19 continues to support FIPS 140-2, and FIPS 140-3 support is available starting in 8.19.15.</strong> That gives you two paths:</p><ul><li><strong>Stay on 8.19.15+, and upgrade the standard in place.</strong> If you're not ready to move to 9.x, you can switch from FIPS 140-2 to 140-3 on your current major version. Your cluster stays put, and your compliance posture moves forward. FIPS 140-2 certificates remain valid through September 2026, so you have a window. But the earlier you transition, the less you're depending on a sunset timeline, and you can also benefit from new features!</li><li><strong>Move to 9.4, and land on 140-3 directly.</strong> If you're planning a major version upgrade anyway, 9.4 gives you FIPS 140-3 from the start, along with everything in the 9.x line: Elasticsearch Query Language (ES|QL) latest functionalities, improved ingest, updated security detections, and performance improvements. No compliance trade-off required.</li></ul><p></p><p></p><p>Stay on 8.19.15+</p><p>Upgrade to 9.4</p><p>FIPS standard</p><p>140-3 (from 8.19.15)</p><p>140-3 (from the start)</p><p>Major version change</p><p>No</p><p>Yes</p><p>New 9.x features</p><p>No</p><p>Yes</p><p>FIPS 140-2 certificates valid until</p><p>September 2026</p><p>N/A (lands on 140-3 directly)</p><p>Kibana covered</p><p>Yes</p><p>Yes</p><p></p><p>Either way, the configuration model will feel familiar. You enable FIPS mode in your YAML config, point to a FIPS-validated JVM, and the stack handles the rest. And Kibana is covered on both paths: Your visualization and analytics layer operates within the same compliant boundary as Elasticsearch.</p><h2>How to enable FIPS 140-3 in Elasticsearch 9.4</h2><p>FIPS 140-3 support is available now in Elastic 9.4 for self-managed deployments with a Platinum or Enterprise subscription, the same licensing tier as FIPS 140-2. For setup instructions, supported JVM versions, configuration details, and known limitations, see the <a href="https://www.elastic.co/docs/deploy-manage/security/fips">FIPS compliance documentation</a>.</p><p><a href="https://www.elastic.co/downloads/elasticsearch"><strong>Download Elastic 9.4</strong></a>, and give your compliance team some good news.</p><p>Questions? Your Elastic account team can help scope a FIPS deployment, or drop into the <a href="https://discuss.elastic.co/">Elastic community forums</a> to compare notes with other operators running in regulated environments.</p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p><h2>Frequently asked questions</h2><p><strong>Does Elasticsearch support FIPS 140-3?</strong></p><p>Yes. FIPS 140-3 support for Elasticsearch and Kibana is generally available in Elastic 9.4 for self-managed deployments. All cryptographic operations are delegated to the Bouncy Castle FIPS Java API 2.0, a FIPS 140-3 validated module. A Platinum or Enterprise subscription is required.</p><p><strong>When does FIPS 140-2 expire?</strong></p><p>NIST stopped accepting new FIPS 140-2 module submissions. Existing FIPS 140-2 certificates remain valid through September 2026. Organizations running Elasticsearch in FIPS 140-2 mode should plan their migration before that deadline to avoid gaps in their compliance documentation.</p><p><strong>Can I migrate from FIPS 140-2 to FIPS 140-3 without upgrading to Elasticsearch 9.x?</strong></p><p>Yes. FIPS 140-3 support is available in Elastic 8.19.15 and later. You can switch from FIPS 140-2 to FIPS 140-3 on your current major version without moving to 9.x. Alternatively, upgrading directly to Elastic 9.4 lands you on FIPS 140-3 from the start.</p><p><strong>What cryptographic changes does FIPS mode make in Elasticsearch?</strong></p><p>When FIPS mode is enabled, Elasticsearch restricts all operations to FIPS-approved cipher suites, replaces bcrypt with PBKDF2 for password hashing, and delegates all cryptographic functions to the Bouncy Castle FIPS provider. Non-approved cipher suites are rejected, not just deprioritized.</p><p><strong>Does Kibana support FIPS 140-3?</strong></p><p>Yes. Kibana operates in a Node.js environment configured with a FIPS-compliant OpenSSL 3 provider. Both Elasticsearch and Kibana run within clearly defined cryptographic boundaries when FIPS mode is enabled.</p><p><strong>Why isn't a FIPS-hardened OS or load balancer sufficient for FedRAMP compliance?</strong></p><p>FedRAMP's network security requirements apply to cryptographic operations at the application layer, not just at the network boundary. Intranode communication, keystore encryption and credential hashing must all occur inside a validated cryptographic module. A compliant perimeter around a non-compliant application layer is an audit finding, not a FIPS deployment.</p><p><strong>Is Elasticsearch FIPS 140-3 support available on Elastic Cloud?</strong></p><p>Elastic 9.4 FIPS 140-3 support covers self-managed deployments. Cloud deployment availability is not covered in this announcement. Contact your Elastic account team or check the Elastic documentation for cloud FIPS roadmap details.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/fips-140-3-elasticsearch-kibana</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/fips-140-3-elasticsearch-kibana</guid>
    <category><![CDATA[Operations]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Fabio Busatto]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1092d0334a5b571c/6a572545f71286c608af57e4/3e67c4eeaa3d9411e97ee8b2ae74078a8177a911-2048x1143.avif" length="0" type="image/*"/>
    <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[How we doubled vector search throughput on Elasticsearch Serverless]]></title>
    <description><![CDATA[How we brought Elasticsearch's native SIMD scoring engine to serverless, and why serverless is where vector search innovation happens next.]]></description>
    <content:encoded><![CDATA[<p>We've brought simdvec, Elasticsearch's native single instruction, multiple data (SIMD) vector scoring engine, to serverless. Search throughput nearly doubled under concurrent load, and p99.9 tail latency dropped from 237 ms to 30 ms. By giving simdvec direct access to the blob cache's memory-mapped regions, serverless now runs the same zero-copy SIMD kernels as stateful, with identical recall and zero heap overhead. And because serverless gives us control over the entire storage layer, we believe it's where vector search will be fastest. Here's how we got there.</p><h2>Vector Search on Elasticsearch Serverless</h2><p><a href="https://www.elastic.co/search-labs/blog/elasticsearch-serverless-stateless-architecture">Elasticsearch Serverless</a> is built on Stateless Elasticsearch, a fully decoupled compute and storage architecture where index data lives in remote object storage and search nodes maintain only a local cache. For vector search to be fast on this architecture, the scoring engine needs to work directly with the local cache, not copy it to the heap first.</p><p>Elasticsearch <a href="https://www.elastic.co/search-labs/blog/elasticsearch-vector-search-simdvec-engine">simdvec</a> is the engine behind every vector distance computation in Elasticsearch. It provides hand-tuned AVX-512 and NEON kernels, bulk scoring with explicit prefetching, and off-heap memory access that keeps data flowing from storage straight to CPU registers. On stateful Elasticsearch, simdvec has always had a direct fuel line: Memory-mapped files feed native pointers straight into SIMD intrinsics. On serverless, the data was sitting right there in the blob cache's memory-mapped regions, in exactly the right form, but there was no path connecting it to the scoring engine.</p><p>We've now built that path. simdvec runs on Serverless with the same off-heap, native SIMD scoring as stateful. And because serverless gives us control over the entire storage layer, this is just the beginning.</p><h2>Premium fuel only: why simdvec requires off-heap memory for vector scoring</h2><p>simdvec's speed comes from working directly with off-heap memory. It takes a native pointer to memory-mapped data and passes it straight to C++ SIMD intrinsics. No intermediate copies, no heap allocations. The data flows from storage straight to CPU registers. This matters more than it sounds: simdvec's kernels process vectors faster than the data can be copied, so any copy in the path becomes the bottleneck, not the scoring itself.</p><p>On stateful Elasticsearch, this just works. Lucene memory-maps index files from local disk, and the scorer extracts a native pointer directly from the mapped region. This is the path that delivers the <a href="https://www.elastic.co/search-labs/blog/elasticsearch-vector-search-simdvec-engine#thousands-at-a-time">benchmark numbers</a> we've published, and it's what we wanted to bring to serverless. To see how, we first need to understand how serverless stores and accesses data.</p><h2>The serverless blob cache: how Elasticsearch stores vector data</h2><p>In the stateless architecture, the primary copy of all index data lives in remote object storage, such as S3. Each search node maintains a local cache (called the <em>blob cache</em>) that keeps recently and frequently accessed portions of the index data on local SSD. The frozen tier on stateful Elasticsearch uses the same architecture: Searchable snapshots are backed by a similar blob cache that memory-maps regions from remote storage onto local disk. When a search hits cached data, it's served from fast local storage. When it misses, the blob cache fetches the data from the remote store and caches it for future queries.</p><p>The blob cache is organized into fixed-size memory-mapped regions, 16MB by default. It manages its own lifecycle: tracking which regions are in use, applying a <a href="https://www.elastic.co/search-labs/blog/searchable-snapshots-benchmark">least-frequently-used eviction policy</a> when the cache is full, and reference counting to ensure regions aren't evicted while being read. The regions are still memory-mapped through the OS, but the blob cache controls which regions exist, which are populated, and when they're reclaimed. On stateful, those decisions are left entirely to the OS.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd17ed30cb6a59275/6a46944fde977731a4ca54b5/e7a1b50ef019d1b5a12d49c7457d63a026e1edd0-727x496.png" alt="Diagram showing data flow between Remote Object Storage and simdvec. The top box labeled “Remote Object Storage” lists S3, GCS, and Azure Blob, with an arrow marked “fetch on miss” pointing to a larger box labeled “Blob Cache.” Inside the Blob Cache are regions numbered 0–5 plus two empty slots, each 16 MB. Regions 0, 1, 3, 4 are green and labeled “cached,” Region 2 is blue and labeled “in use,” Region 5 is yellow and labeled “evicting,” and two gray boxes are labeled “empty.” A legend explains the color codes. A downward arrow labeled “direct memory” connects Blob Cache to a dark box labeled “simdvec – native SIMD scoring.”" /><p>Crucially, because each region is memory-mapped, the blob cache already holds vector data in exactly the form simdvec needs. But before <a href="https://github.com/elastic/elasticsearch/pull/141718">we built the connection</a>, there was no way to get at it. Every vector comparison was copied into a heap array and handed to a slower scorer. No direct memory pointers, no SIMD, and garbage collection pressure on every call.</p><h2>Unified scoring: one SIMD path for all storage tiers</h2><p>We introduced a new abstraction that lets the scorer safely borrow direct memory from whatever storage layer is underneath, just long enough to run the SIMD computation. If the data is available as direct memory, simdvec's native kernels run. If not (data not yet cached or spanning a region boundary), the scorer falls back to a heap copy. In practice, the fallback is rare.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltadef4930fe865721/6a469452a65a6b4e1dbeff0f/fd9583ffce724665018af46ce0409dc4e0825078-828x259.png" alt="Side‑by‑side comparison diagram labeled “Before” and “After.” The “Before” section shows four boxes: blue “Stateful – local mmap,” green “simdvec – native SIMD ✓,” yellow “Serverless – blob cache,” and red “Java scorer – no SIMD ✗.” Arrows indicate a green “direct ptr” from Stateful to simdvec and a red “heap copy” from Serverless to Java scorer, with the caption “two paths, two implementations.” The “After” section shows three boxes: blue “Stateful – local mmap,” yellow “Serverless – blob cache,” and green “simdvec – native SIMD ✓,” with two green arrows labeled “direct” pointing to simdvec and the caption “one engine, one code path, all tiers.&quot;" /><p>This gave us a single scoring entry point across all tiers:</p><ol><li><p><strong>Stateful</strong> (local disk): The scorer extracts a native pointer from the OS memory map.</p></li><li><p><strong>Blob cache</strong> (serverless, frozen tier): The scorer borrows a direct memory slice from a cache region.</p></li><li><p><strong>Fallback</strong>: The scorer copies bytes to the heap. Rare in practice.</p></li></ol><p>The scorer doesn't know which tier it's running on, and it doesn't need to. It also means we no longer maintain separate scoring implementations; previously, there was a fast native path for stateful and a slower path for everything else. Now every improvement to simdvec benefits all tiers automatically, including its most powerful capability: bulk scoring.</p><h2>Bulk vector scoring across blob cache regions</h2><p>A single query may score thousands of candidate vectors. simdvec's <a href="https://www.elastic.co/search-labs/blog/elasticsearch-vector-search-simdvec-engine#thousands-at-a-time">bulk scoring</a> processes these in batches with multi-accumulator inner loops, query amortization, and cache-line prefetching, up to 4x faster than single-vector alternatives when data exceeds CPU cache.</p><p>Search over an Inverted file (IVF) index is where bulk scoring has the most impact. The query selects a set of candidate posting lists and sweeps through the quantized vectors, scoring them in large batches against the query vector. On stateful, those vectors live in one contiguous memory-mapped file, so bulk scoring resolves them with straightforward pointer arithmetic and scores a batch in a single native call.</p><p>On serverless, a sweep through a posting list may cross blob cache region boundaries. We extended the direct memory abstraction with a bulk access method that resolves multiple vector offsets to their respective cache regions in a single call. If all vectors in the batch are cached and none cross a region boundary, the scorer gets a direct memory slice and passes the whole batch to simdvec's native bulk kernel with the same prefetching and pipelining as stateful. When a vector does cross a boundary, the system falls back to per-vector scoring: still zero-copy, just without the batching benefit. With 16MB regions and 1024-byte vectors, that happens roughly once every 16,000 vectors.</p><p>simdvec's bulk scoring architecture, the key differentiator highlighted in the simdvec <a href="https://www.elastic.co/search-labs/blog/elasticsearch-vector-search-simdvec-engine">benchmarks</a>, now operates on serverless with the same characteristics that make it fast on stateful. So how does it perform in practice?</p><h2>simdvec on Elasticsearch Serverless: vector search lap times</h2><p>We benchmarked with an 18 million vector <a href="https://github.com/elastic/rally-tracks/tree/master/msmarco-v2-vector">MSMARCO</a> dataset at 1024 dimensions, using IVF with Better Binary Quantization (BBQ) 1-bit quantization. All results are on a warm blob cache with the full dataset resident in local cache regions, so we're measuring the scoring path rather than remote fetch latency.</p><p><strong>Throughput.</strong> Under concurrent load, search throughput nearly doubled, jumping from 398 to 739 ops/s. Single-client gains were 23-39%, but the real difference shows up under concurrency: The improvement was 2-3x larger because eliminating heap copies removes the GC pressure and allocation contention that previously throttled concurrent scoring.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt277d31993bac780c/6a4694559450737b7530d222/fc152731c4b99b85a448e8bd4e01915fefcb55a3-919x533.png" alt="Bar chart titled “Search Throughput — Baseline vs Zero‑Copy (Median ops/s).” It compares median throughput between Baseline (heap‑copy) and Zero‑Copy (DirectAccessInput) across eight knn configurations. Each group shows a gray Baseline bar and a taller green Zero‑Copy bar with percentage improvements labeled above. The y‑axis shows median throughput in operations per second, ranging up to 900. The subhead notes that percentage labels indicate improvement." /><p><strong>Tail latency.</strong> The direct memory path transformed tail latency under load:</p><ul><li><p><em>p99.9</em> dropped from 237 ms to 30 ms (87% reduction).</p></li><li><p><em>p99.99</em> dropped from 9.1 seconds to 55 ms (99.4% reduction).</p></li></ul><p><em>p100</em> dropped from 11.4 seconds to under 100 ms.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6002a4d4bf8cbd8f/6a469458c71ec49c21b98453/003874d85261507d95274504a83bc016af0beb13-818x555.png" alt="Line graph titled “Tail Latency Collapse — knn‑10‑10 Multi‑Client.” The chart compares Baseline (heap‑copy) and Zero‑Copy (DirectAccessInput) latency across percentiles p50 to p100 on a logarithmic scale. The red Baseline line rises sharply, while the green Zero‑Copy line remains low. Labels mark key points. The caption notes Rally benchmark details and log scale." /><p>The worst-case outliers that previously took seconds now complete in tens of milliseconds. The heap-copy-induced queueing that caused latency spikes is gone.</p><p>Recall is identical. The same vectors are scored, producing the same results. And we're just getting started.</p><h2>Beyond parity: what Elasticsearch Serverless can do for vector search that stateful can't</h2><p>Reaching parity with stateful was the goal. But the more interesting realization is what the stateless architecture lets us do that stateful can’t.</p><p>On stateful, the OS controls memory-mapped file behavior: which pages stay resident, when to evict, how aggressively to read ahead. The application can offer hints, but they apply to entire file mappings, and the kernel may ignore them. Worse, search and indexing happen concurrently on the same node, so a hint that benefits one access pattern can hurt another. In practice, to balance different needs, you have to be conservative.</p><p>On serverless, two things are fundamentally different. The blob cache manages its own memory-mapped regions with full application-level control. And serverless <a href="https://github.com/elastic/elasticsearch/issues/147626">separates indexing and search onto dedicated tiers</a>: Search nodes never merge, indexing nodes never serve queries. No conflicting access patterns means we can be aggressive with memory advice. Here’s what we’re working on:</p><ul><li><p><strong>Per-region memory advice.</strong> The blob cache knows what type of data each region holds. It can issue <a href="https://github.com/elastic/elasticsearch/issues/147625">random-access hints for rescoring regions</a>, where raw float32 vectors are read in unpredictable order and the kernel’s default readahead would waste memory on pages that will never be used. It can apply sequential readahead for scans through quantized vectors. On the indexing tier, merges read data sequentially, so aggressive readahead brings pages in before they're needed, with no risk of harming concurrent random reads that simply aren't happening on that node.</p></li><li><p><strong>Cache-aware prefetching.</strong> simdvec already prefetches at the CPU cache-line level. On serverless, we can coordinate this with the blob cache's knowledge of region residency, prefetching at multiple levels: remote store to cache, OS pages to RAM, and cache lines to CPU. The blob cache can <a href="https://github.com/elastic/elasticsearch/pull/147964">tell the scorer</a> which regions are resident before scoring begins, avoiding work on data that would trigger a remote fetch.</p></li><li><p><strong>Workload-aware eviction.</strong> The blob cache can prioritize retaining data that vector search depends on: IVF centroid indexes that are checked on every query or quantized vectors that are scored in bulk, over data that's accessed infrequently. The OS page cache evicts based on generic heuristics with no understanding of what the data represents. On serverless, eviction policy can be tuned to the workload.</p></li></ul><p>The blob cache gives us a level of control over the memory hierarchy that the OS page cache simply can’t. This is why we see serverless as the most promising platform for the next generation of vector search performance work. Not just matching stateful, but surpassing it. And vectors are just the beginning.</p><h2>Vector search on Elasticsearch Serverless: what we shipped and what's next</h2><p>simdvec now runs everywhere Elasticsearch runs (stateful, serverless, and frozen tier) with the same native SIMD scoring, the same bulk scoring, and the same off-heap efficiency. The abstraction we built is general-purpose and already wired through every layer in the storage chain, so the same approach could benefit term lookups, aggregations, sorting, and stored field retrieval in the future.</p><p>Elasticsearch Serverless is where we're investing most heavily in vector search performance. Every improvement to simdvec, every optimization to the blob cache, and every new storage-level improvement lands here first. If you're choosing where to run your vector workloads, serverless is the platform that keeps getting faster. You can get started with a free <a href="https://cloud.elastic.co/registration">Elastic Cloud trial</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/vector-search-serverless-simdvec-throughput</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/vector-search-serverless-simdvec-throughput</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Chris Hegarty,Lorenzo Dematte]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd17ed30cb6a59275/6a46944fde977731a4ca54b5/e7a1b50ef019d1b5a12d49c7457d63a026e1edd0-727x496.png" length="0" type="image/png"/>
    <pubDate>Thu, 28 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[How Elasticsearch cuts time-series storage by 34% with synthetic _id and bloom filters]]></title>
    <description><![CDATA[Learn how synthetic _id uses bloom filters to cut time-series storage by 34% while maintaining full API compatibility.]]></description>
    <content:encoded><![CDATA[<p>Synthetic <code>_id</code> reduces time-series index storage by up to 34% and eliminates 6% CPU overhead at ingest. Instead of building an inverted index for <code>_id</code>, Elasticsearch computes the document identifier on the fly from <code>_tsid</code> and <code>@timestamp</code>, using a bloom filter for deduplication. This optimization ships in Elasticsearch 9.4 and is already live on Elastic Cloud Serverless.</p><p>This post is a deep dive into the implementation. For context on how synthetic <code>_id</code> fits into the broader metrics performance story, see <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-metrics-engine-30x-faster-prometheus">How we rebuilt Elasticsearch as a leading columnar metrics datastore</a> to achieve up to 6.6x improvement in storage efficiency and 50% improvement in indexing throughput for OpenTelemetry metrics.</p><p>We'll start by explaining why the <code>_id</code> field is expensive for time-series workloads. We'll then describe how synthetic <code>_id</code> works and how it uses a bloom filter to optimize document deduplications instead of maintaining a traditional inverted index. Finally, we'll share the performance results from our benchmarks and serverless production deployments.</p><h2>The hidden cost of _id in time-series indices</h2><p>Time-series indices are a specialized index mode optimized for metrics, logs, traces, and other timestamped data. They store sequences of data points (like CPU usage, stock prices, or sensor readings) that track changes to specific entities over time. In Elasticsearch, each of these data points is indexed as a document with a unique identifier called <code>_id</code>. This identifier is used to look up, update, or delete specific documents. When a document is indexed in Elasticsearch, the system checks whether a document with the same <code>_id</code> already exists. Depending on the operation type (<code>op_type</code>), an existing document is either replaced (<code>index</code>) or the new document is rejected (<code>create</code>); the latter is the most common path for metrics ingestion.</p><p>To perform this lookup efficiently, Elasticsearch builds an <a href="https://en.wikipedia.org/wiki/Inverted_index">inverted index</a> for the <code>_id</code> field. This inverted index maps each <code>_id</code> value to its location in the index, enabling fast document lookups. Until version 8.11, the <code>_id</code> value was also stored separately in order to be returned in search results and other APIs. From 8.11 and onwards, we optimized Elasticsearch to only store this value temporarily for document replication purposes, the value being quickly merged away and reconstructed on demand.</p><p>For many use cases, building the inverted index and storing it is an acceptable overhead. But for time-series data, like metrics or traces, the cost can add up quickly. Our experiments showed that building the inverted index for the field <code>_id</code> adds 6% CPU overhead compared to indexing without it. In some extreme cases, we benchmarked that it could reduce indexing throughput by 25%.</p><p>This overhead is especially painful for time-series workloads where data points are typically small (often just a timestamp and a few numeric values) and compress extremely well. The <code>_id</code> field, however, doesn't benefit from the same compression. As a result, the inverted index for <code>_id</code> can represent a disproportionate share of the total storage. In our benchmarks with OpenTelemetry (OTel) metrics, the <code>_id</code> inverted index alone consumed around 5 bytes of the total 25 bytes per data point.</p><p>We considered several approaches to eliminate this overhead:</p><ul><li><p>Stop indexing <code>_id</code> and checking for duplicates: This would be the simplest solution, but without deduplication, duplicate data points could corrupt aggregations. A gauge average, for instance, would be skewed by repeated values.</p></li><li><p>Accept duplicates during indexing, deduplicate at query time: This preserves correctness but adds overhead to every query, degrading dashboard responsiveness.</p></li><li><p>Deduplicate during segment merges: Duplicates would eventually be removed, but queries on unmerged segments would still return results with duplicates.</p></li><li><p>Synthetic <code>_id</code>: Compute the document identifier on the fly from fields that already uniquely identify each data point, and use a lightweight bloom filter for deduplication instead of a full inverted index.</p></li></ul><p>We chose synthetic <code>_id</code> because it maintains correctness at ingest time while eliminating the storage and CPU overhead of the traditional approach. And we decided to implement it for time-series indices because they’re very well suited for this optimization.</p><p>In time-series indices, the <code>_id</code> isn’t arbitrary. Each document has a <strong>time series identifier</strong> (<code>_tsid</code>) and a <strong>timestamp</strong> (<code>@timestamp</code>). The <code>_tsid</code> is generated from the <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/time-series-data-stream-tsds#time-series-dimension">dimensions fields</a> of the document (like <code>host.name</code>, <code>pod.name</code>, or <code>sensor_id</code>), while the <code>@timestamp</code> marks the point in time of the document. Together, these two fields uniquely identify the document: There can only be one data point for a given time series at a given moment in time. This means we can derive the <code>_id</code> from the <code>_tsid</code> and <code>@timestamp</code> field values, rather than storing it separately.</p><h2>How does synthetic _id work in Elasticsearch?</h2><p>With synthetic <code>_id</code>, Elasticsearch computes the document identifier on the fly as the combination of the <code>_tsid</code> and <code>@timestamp</code> fields. This computed value is used wherever <code>_id</code> would normally be used: in API responses, for document lookups, and for deduplication. However, it’s never stored in an inverted index nor is it stored on disk for later retrieval.</p><p>The challenge is deduplication. When a new document arrives, Elasticsearch must verify that no document with the same <code>_id</code> already exists. Without an inverted index on <code>_id</code>, how can we perform this check efficiently?</p><h3>How synthetic _id simulates an inverted index without building one</h3><p>Our Elastic Lucene experts suggested a clever idea: Since <code>_tsid</code> and <code>@timestamp</code> are already stored as doc values, we could expose our own custom Lucene postings format that simulates an inverted index without actually building one.</p><p>This means that when Elasticsearch needs to look up a document by its <code>_id</code>, it uses the same code path as usual: It queries the underlying Lucene index to look up the <code>_id</code> term. But instead of hitting a real inverted index, our custom postings format intercepts the call, extracts the <code>_tsid</code> and <code>@timestamp</code> encoded in the synthetic <code>_id</code>, and uses their doc values to locate the document. Because time-series indices are sorted by these fields, documents belonging to the same time series are stored contiguously. This allows Elasticsearch to skip large subsets of nonmatching documents (sometimes entire segments) to find the target document(s) quickly.</p><p>While this process is efficient, it can involve several random-access reads: looking up the <code>_tsid</code> value, scanning for matching documents, and reading timestamps. For the common case in time-series indices where we don’t expect the document to already exist, we wanted to fail fast without touching doc values at all.</p><h3>Bloom filters for fast membership testing</h3><p>We solve this problem using a <a href="https://en.wikipedia.org/wiki/Bloom_filter"><strong>bloom filter</strong></a>, a probabilistic data structure that can quickly answer the question <em>Could this element be in the set?</em> with a small risk of false positives but no risk of false negatives. In other words, a bloom filter might occasionally say <em>yes</em> when the answer is actually <em>no</em>, but it will never say <em>no</em> when the answer is <em>yes</em>.</p><p>When a document is indexed, its synthetic <code>_id</code> is added to the bloom filter. When a new document arrives, we first check the bloom filter. If the bloom filter says <em>no</em>, we know for certain that no document with this <code>_id</code> exists and we can proceed with indexing immediately. If the bloom filter says <em>maybe yes</em>, we fall back to the more expensive verification using the <code>_tsid</code> and <code>@timestamp</code> doc values.</p><h3>Synthetic _id indexing workflow: step by step</h3><p>Let's walk through what happens when a document is indexed into a time-series index with synthetic <code>_id</code> enabled:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt387a8573301f120e/6a3e41e975bd4076e6a77b5d/61ef279f09c5447009d3695f154129fba6fe510d-1048x1462.png" alt="Flowchart on a dark background showing document indexing steps using synthetic IDs, bloom filters, and duplicate handling paths." /><ol><li><p><strong>Compute the synthetic </strong><strong><code>_id</code></strong>: Elasticsearch calculates <code>_id</code> as a combination of <code>_tsid || @timestamp</code>.</p></li><li><p><strong>Check the live version map</strong>: Like today, we first check an in-memory map of recently indexed documents. If the document is present in this map, we can handle the duplicate immediately.</p></li><li><p><strong>Filter segments by timestamp</strong>: Time-series indices are sorted by <code>_tsid</code> and <code>@timestamp</code>. We can skip any segment whose timestamp range does not overlap with the incoming document's timestamp.</p></li><li><p><strong>Check the bloom filter</strong>: For each candidate segment, we test whether the <code>_id</code> might exist using the bloom filter.</p></li><li><p><strong>Verify if needed</strong>: If the bloom filter returns a positive result, we look up the document using the <code>_tsid</code> and <code>@timestamp</code> doc values. Since documents are sorted by these fields, this lookup is efficient.</p></li><li><p><strong>Index the document</strong>: If no existing version is found, the document is indexed. The <code>_id</code> is added to the segment's bloom filter, but no inverted index is built and the field value is never stored.</p></li></ol><p>In the common case where new data arrives with recent timestamps, step 3 eliminates most segments from consideration, and step 4 quickly confirms that the document is new. The expensive verification in step 5 only happens on bloom filter false positives, which are expected to be rare.</p><h3>Bloom filter false positive rate: how Elasticsearch keeps it low</h3><p>One challenge with bloom-filter-based deduplication is controlling the false positive rate without sacrificing the storage efficiency we were after. To size bloom filters effectively, we consider the number of data points in each segment and target both a low false positive rate and a bit set saturation below 50%.</p><p>The saturation target serves a specific purpose: When segments are merged, we OR the bit sets rather than rebuilding bloom filters from scratch. This makes merges fast but means the false positive rate converges toward 100% as segments are merged repeatedly. Keeping saturation below 50% before merging buys headroom, delaying that convergence.</p><p>The low false positive rate target is justified by access patterns: Recent segments are checked far more often than older ones, since we prune the search space based on data point timestamps. Older, heavily merged segments with degraded bloom filters are unlikely to be checked.</p><h2>Synthetic _id performance benchmarks: indexing and storage</h2><p>We ran extensive benchmarks to validate our implementation.</p><h3>Indexing throughput</h3><p>A core goal of this effort was to match or improve on existing indexing throughput. In principle, the new approach does less work: Building an inverted index for <code>_id</code> requires hashing each value, building and maintaining complex data structures in memory, and flushing them to disk. These structures must also be reconstructed during segment merges, adding CPU and I/O overhead in high-throughput use cases.</p><p>Building a bloom filter isn't free (we still hash each value), but the memory footprint is smaller and there are no complex data structures to maintain or flush. The bloom filter is also cheap to merge: When possible, we simply OR the bit sets together rather than rebuilding from scratch.</p><p>The main cost of synthetic <code>_id</code> comes from verifying potential duplicates using doc values. However, this cost is mitigated by two factors: First, bloom filter false positives are rare, so most documents skip this step entirely. Second, time-series indices are sorted by <code>_tsid</code> and <code>@timestamp</code>, which means doc value lookups can skip large blocks of nonmatching documents efficiently.</p><p>In practice, that's exactly what we observed. Even accounting for the extra seeks needed to verify matches against the tsid and timestamp when a bloom filter returns a positive, throughput came out comparable or better than before. The savings from not building and merging the inverted index outweigh the occasional cost of a false positive check, as confirmed by our <a href="https://elasticsearch-benchmark-analytics.elastic.co/app/dashboards#/view/f7e091a0-1db1-11ed-920a-3b1141502d24?_g=(refreshInterval:(pause:!t,value:60000),time:(from:'2026-03-16T00:00:00.000Z',to:'2026-03-19T23:30:00.000Z'))&amp;_a=(viewMode:view)">nightly benchmarks</a>:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte73bd65d8545658d/6a3e41eca48ab9170737c817/32b5b068ef6dd7cd512f12dc9d50d8769f2930d4-1999x509.png" alt="Line graph titled “nightly‑tsdb‑indexing‑throughput,” showing nightly benchmark results for document indexing rates in docs per second over three days, with four colored lines representing different indexing operations." /><h3>Storage savings</h3><p>In our benchmarks with OTel metrics, synthetic <code>_id</code> reduced storage by approximately 5 bytes per data point. For a dataset where documents average 25 bytes per data point, this represents a 20% reduction in storage from this single optimization alone.</p><p>These results were soon confirmed by our <a href="https://elasticsearch-benchmarks.elastic.co/#tracks/tsdb/nightly/default/90d">nightly benchmarks</a>.The chart below shows the storage footprint reduction over time as we enabled the synthetic <code>_id</code> feature on March 19, 2026:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3cb1cf931ecebba9/6a3e41efb8c8ed6dec833e9d/0c0ac1f2c9716a6f99b6e055535db0def77bf930-1898x956.png" alt="Line graph titled “Disk usage” showing TSDB and downsampling data from March 15 to March 23, 2026, with disk usage measured in gigabytes per 24 hours. A teal line for TSDB drops near March 19 and stabilizes around 1.9 GB, while a brown line for downsampling decreases to 2.3 GB after the same date." /><p>Our standard time series database (TSDB) benchmark showed a reduction from 2.5 GiB to 1.9 GiB (24%). Similarly the time-series downsampling benchmark showed a comparable reduction from 3 GiB to 2.3 GiB (23%).</p><p>Another benchmark, more focused on metrics, <a href="https://elasticsearch-benchmark-analytics.elastic.co/app/dashboards#/view/37270832-cd2d-4ea7-8222-e61e8ad742a3?_g=(refreshInterval:(pause:!t,value:60000),time:(from:'2026-03-16T00:00:00.000Z',to:'2026-03-19T23:30:00.000Z'))&amp;_a=(viewMode:view)">showed an even better reduction</a>, from 3.0 GiB to 2.0 GiB (34%):</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd66d7617467e01c0/6a3e41f2809064d8e0e00d7f/2a243db0f3ad95e761fb334ffd6fea47f7787eee-1999x447.png" alt="Line graph titled “Dataset size,” showing two turquoise lines that decline from March 16 to 19, 2026, each representing a different TSDB metric. The x‑axis marks daily timestamps, and the y‑axis shows dataset size decreasing." /><h2>API compatibility</h2><p>An important design goal was maintaining compatibility with existing Elasticsearch APIs. With synthetic <code>_id</code>, all document APIs continue to work as expected: Bulk, Get, Update, Delete, Reindex, and Update/Delete by Query. This compatibility layer also limited the blast radius of the change, ensuring any issues would be contained to the internal implementation.</p><p>When the <code>_id</code> isn’t provided in an API request, Elasticsearch computes it from the <code>_tsid</code> and <code>@timestamp</code> fields. To check if the document already exists, it first queries the bloom filter and, if needed, falls back to doc values. The <code>_id</code> is also synthesized on demand from doc values when returning documents in search results or API responses.</p><p>One case that requires special handling is searching or filtering by <code>_id</code> prefix or pattern. Such queries require scanning many documents to find matching documents, and while this works correctly, it incurs a performance penalty compared to a direct <code>_id</code> lookup. We don’t expect this use case to be common for time-series indices though.</p><h2>Elasticsearch 9.4 and Elastic Cloud Serverless availability</h2><p>The synthetic <code>_id</code> feature will be released in Elasticsearch 9.4.0 and is already available on <a href="https://www.elastic.co/cloud/serverless">Elastic Cloud Serverless</a>.</p><p>No configuration is required: The feature is enabled by default, and newly created time-series indices (including those created on datastream rollover) will automatically benefit from this optimization. Existing time-series indices created before 9.4 will continue to create inverted indices for the <code>_id</code> field.</p><p>We expect synthetic <code>_id</code> to perform well across all time-series use cases. However, in some very specific, update-heavy use cases, if you encounter performance issues, the feature can be disabled by setting <code>index.mapping.synthetic_id</code> to <code>false</code> for new indices.</p><h2>Summary: synthetic _id storage and performance gains</h2><p>In this article, we’ve presented how synthetic <code>_id</code> eliminates the storage and compute overhead of document identifiers in time-series indices. By computing <code>_id</code> on the fly from <code>_tsid</code> and <code>@timestamp</code>, and using a bloom filter for deduplication, we achieve comparable or better indexing performance with up to 34% reduction in storage footprint while maintaining full API compatibility. For users running large-scale time-series workloads, this translates directly into lower infrastructure costs.</p><h2>Roadmap: what comes after synthetic _id</h2><p>Synthetic <code>_id</code> is part of a broader effort to reduce storage overhead in Elasticsearch.</p><ul><li><p><strong>Sequence number trimming:</strong> Every document carries a sequence number for replication and concurrency control. For append-only time-series data, these become redundant after segments are merged. Elasticsearch 9.4 now trims them during merges to reclaim even more storage: We'll cover this optimization in detail in an upcoming blog post.</p></li><li><p><strong>Synthetic _id beyond time-series:</strong> We’re exploring how to bring synthetic <code>_id</code> to regular indices by letting users declare which fields uniquely identify their documents and configuring index sorting on those fields to enable efficient lookups.</p></li></ul><p>Stay tuned!</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-synthetic-id-time-series-storage</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-synthetic-id-time-series-storage</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Tanguy Leroux,Francisco Fernández Castaño,Anton Persson]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbe7de1872f1527e3/6a17de303e9e452974ba1374/a70c5403064d5bbceff66a17373332362227f13c-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 28 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Does MCP make search obsolete? Not even close]]></title>
    <description><![CDATA[Explore why search engines and indexed search remain the foundation for scalable, accurate, enterprise-grade AI, even in the age of MCP, federated search, and large context windows.]]></description>
    <content:encoded><![CDATA[<p>With the rise of large language models (LLMs), agent frameworks, and new protocols like Model Context Protocol (MCP), a provocative question is starting to surface:</p><strong>Do we still need a search engine at all?</strong><p>If agents can call tools on demand and models can reason over massive context windows, why not just fetch data live from every system and let the LLM figure it out?</p><p>It’s a reasonable question. It’s also the wrong conclusion.</p><p>The reality is that MCP and agent tooling don’t eliminate the need for search. They make the quality of search <strong>more critical than ever</strong>. In this blog, we’ll explore why MCP, federated search, and large context windows don’t replace search engines and why indexes remain the foundational layer for scalable, accurate, enterprise-grade AI.</p><h2><strong>What MCP actually is (and what it is not)</strong></h2><p>MCP is a <strong>coordination protocol</strong>. It standardizes how an agent requests information or actions from external systems.</p><p>What MCP <em>doesn’t</em> do:</p><ul><li><p>Rank results across systems.</p></li><li><p>Understand relevance across heterogeneous data.</p></li><li><p>Normalize schemas or metadata.</p></li><li><p>Data transformations or enrichments at scale.</p></li><li><p>Apply consistent security and permissions.</p></li><li><p>Optimize for latency, cost, or scale.</p></li></ul><p>In other words, <strong>MCP tells agents </strong><em><strong>how</strong></em><strong> to ask for data, not </strong><em><strong>which</strong></em><strong> data matters most</strong>.</p><h2><strong>Modern retrieval requires query intelligence, not just data access</strong></h2><p>In modern enterprise search architectures, retrieval quality is determined long before a query reaches an index. Raw queries — especially those generated by agents — may be incomplete, overly literal, schema-driven rather than intent-driven, and at times syntactically invalid.</p><p>This is why mature search platforms introduce a query intelligence layer that performs query rewriting, entity normalization, synonym expansion, and intent disambiguation before retrieval even begins.</p><p>For example, an agent-generated request such as: “Show severity 2 authentication failures from last sprint” may be rewritten to include authentication synonyms (login, SSO, OAuth), normalized severity mappings, and sprint-to-date-range translation. The result is not just more matches — it is more <em>relevant</em> matches.</p><p>In enterprise AI, retrieval is not a single step. It is a controlled pipeline.</p><p>This distinction is crucial because once MCP-based agents start pulling information live from multiple tools, they recreate a familiar pattern under a new name: <strong>federated search</strong>.</p><h2><strong>MCP-based retrieval is federated search in disguise</strong></h2><p>Federated search isn’t new. Enterprises have tried it for decades.</p><p>The model is simple:</p><ul><li><p>Send the user’s query to multiple systems in parallel (SharePoint, GitHub, Jira, customer relationship management [CRM]).</p></li><li><p>Collect the responses.</p></li><li><p>Merge and present the results.</p></li></ul><p>MCP-driven tool calls follow the same pattern, except that the caller is now an agent instead of a user interface.</p><p>And the same problems resurface.</p><h2><strong>Why federated search breaks down at enterprise scale</strong></h2><ul><li><p><strong>Latency becomes unpredictable:</strong> A federated query is only as fast as its slowest system. Enterprise systems can have wildly different response times and rate limits, so federated queries tend to be <strong>slow and jittery</strong>. Agents must wait for multiple round trips before reasoning can even begin. The result is a laggy experience and unpredictable wait times.</p></li><li><p><strong>Relevance is fragmented:</strong> Because each system ranks results on its own, there’s no unified relevance model. Federated search <strong>cannot apply a single ranking or semantic understanding across all content</strong>, so results often seem disjointed or incomplete. Agents may retrieve <em>correct</em> information but not the <em>most useful</em> information.</p></li><li><p><strong>Context is shallow and incomplete: </strong>Federated systems typically expose only what’s directly accessible through an API call.They rarely surface:</p><ul><li><p>Usage signals, like clicks, dwell time, recency of access, popularity, or authority.</p></li><li><p>Relationships between documents across different systems to correlate the insights.</p></li><li><p>Organizational knowledge beyond a single silo.

This strips agents of the broader context required for high-quality reasoning.
</p></li></ul></li><li><p><strong>Limited filtering and features:</strong> In a federated setup, you can only filter on fields that every system supports (the “lowest common denominator”). If one system doesn’t support a particular filter or facet, you lose that functionality entirely. This severely limits rich search features, like date ranges, categories, or tags.</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6f82101d3ef2019b/6a170ca96f7f04f0219148ac/25bb778f4da9a3cb4f0d4e10af66221b8af73900-1376x768.jpg" alt="Federated search workflow" /><h2><strong>The power of an indexed search</strong></h2><p>Search engines achieve millisecond-level retrieval at massive scale by using specialized data structures, including inverted indexes for lexical search and k‑dimensional trees (k-d trees) for vector-based retrieval. The approach is to <strong>crawl or ingest every source into search engines</strong>, creating a central place of company knowledge. This brings big advantages:</p><ul><li><p><strong>Speed by design:</strong> Searching an index is lightning fast. Queries hit inverted indexes and specialized data structures, avoiding the need to poll each backend system.</p></li><li><p><strong>Relevance that compounds over time:</strong> Search engines that support <strong>semantic search </strong>are capable of comprehending the intent, and machine learning models can rerank results for enterprise contexts. In one Elastic <a href="https://www.elastic.co/blog/elastic-generative-ai-experiences?">experiment</a>, Elastic users see more accurate results when combining vector search with a question-answering (QA) model to extract answers. It gives better precision than keyword matching.</p></li><li><p><strong>Advanced features:</strong> Elastic’s <a href="https://www.elastic.co/search-labs/blog/rag-graph-traversal#:~:text=Retrieval,for%20deeper%2C%20more%20contextual%20retrieval">Graph retrieval augmented generation (RAG) solution</a> shows how structuring an index as a knowledge graph can power more contextual retrieval. In other words, indexes aren’t just backward-looking dumps of text; they can also encode relationships and ontologies that let AI connect the dots across documents.</p></li><li><p><strong>Permission-aware search:</strong> Enterprise AI cannot compromise on security. Indexed search allows:</p><ul><li><p><a href="https://www.elastic.co/docs/reference/search-connectors/document-level-security">Document-level security.</a></p></li><li><p><a href="https://www.elastic.co/docs/deploy-manage/users-roles/cluster-or-deployment-auth/user-roles#roles">Role-based access control.</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/rag-and-rbac-integration">Permission-aware retrieval for RAG and agents.</a></p></li></ul></li></ul><p>Agents see only what users are allowed to see, without leaking data into model prompts or training. Elasticsearch is suitable for the indexed search layer in the diagram below, as it provides the essential components for context engineering.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb88e668cb4814bee/6a170cab8b73cb5d6918a090/8785e7806616273d086a90b3540273fb26d045ae-1392x768.jpg" alt="Essential components of context engineering, highlighting the Elasticsearch role in the indexed search layer." /><h2><strong>Retrieval consistency through search templates and governed execution</strong></h2><p>At scale, retrieval must be predictable, secure, and repeatable. This is where <a href="https://www.elastic.co/docs/solutions/search/search-templates">search templates</a> become critical.</p><p>Search templates act as retrieval contracts between applications, agents, and the search platform. Instead of dynamically constructing queries at runtime, agents invoke pre-defined retrieval patterns that enforce:</p><ul><li><p>Consistent relevance logic</p></li><li><p>Mandatory security filters</p></li><li><p>Cost and latency guardrails</p></li><li><p>Business-specific ranking rules</p></li><li><p>Explicit index and field scope boundaries</p></li></ul><p>In MCP-driven architectures, this becomes even more important. Agents should not dynamically invent retrieval strategies. Instead, MCP tool calls can map directly to approved search templates, ensuring that every retrieval request adheres to enterprise relevance and governance standards.</p><p>This approach shifts retrieval from ad-hoc query execution to controlled retrieval orchestration.</p><h2><strong>Retrieval is now a multi-layer engineering discipline</strong></h2><p>Modern enterprise retrieval is no longer a simple query-to-index operation. It typically includes multiple coordinated layers:</p><ul><li><p>Query understanding — rewriting, expansion, entity resolution</p></li><li><p>Retrieval strategy selection — hybrid search, vector search, graph retrieval, or synthetic query techniques such as Hypothetical Document Embeddings (HyDE), where the system generates a representative answer or expanded context first and retrieves documents using that richer semantic signal.</p></li><li><p>Execution governance — templates, security enforcement, and performance guardrails</p></li><li><p>Ranking and re-ranking — blending lexical precision, semantic similarity, and interaction-derived relevance signals such as click-through patterns, dwell time, and document usage frequency.</p></li></ul><p>When these layers are implemented upstream, agents receive clean, high-confidence context rather than raw, fragmented data.</p><p>This is what makes large-scale agent systems reliable in production environments.</p><h2><strong>Advanced retrieval techniques improve context quality before reasoning begins</strong></h2><p>Modern retrieval systems increasingly use AI-assisted techniques to improve recall and semantic coverage before ranking is applied.</p><p>One example is <a href="https://medium.com/@nirdiamant21/hyde-exploring-hypothetical-document-embeddings-for-ai-retrieval-cc5e5ac085a6">Hypothetical Document Embeddings (HyDE)</a>. Instead of embedding only the original query, the system first generates a hypothetical answer or expanded context, embeds that representation, and retrieves documents based on that richer semantic signal.</p><p>This is particularly useful in enterprise environments where:</p><ul><li><p>Users or agents may not know the exact terminology</p></li><li><p>Knowledge is distributed across silos</p></li><li><p>Important context is implied rather than explicitly stated</p></li></ul><p>Techniques like HyDE improve the probability that relevant documents are retrieved even when the original query is underspecified.</p><p>This reinforces a key principle of enterprise AI: better context retrieval produces better reasoning outcomes.</p><h2><strong>Agents aren’t data engineers; they’re reasoning systems</strong></h2><p>They shouldn’t be responsible for stitching together raw data, reconciling schemas, or compensating for poor retrieval.</p><p>This is where a search platform such as <strong>Elasticsearch</strong> becomes foundational.</p><p>By ingesting data once and normalizing it upstream (through pipelines, mappings, enrichment processors, and prebuilt indexes), Elasticsearch resolves schema mismatches, joins signals across sources, and materializes retrieval-ready views of the data. At query time, the agent receives clean, ranked, semantically enriched results rather than fragmented raw records.</p><p>For example, instead of an agent pulling independently from CRM, ticketing, and documentation systems and attempting to reconcile customer IDs, timestamps, and formats in real time, Elasticsearch can pre-index these sources into a unified customer interaction index with hybrid (keyword + vector) search and relevance ranking. The agent then queries a single, coherent interface and immediately reasons over the most relevant context.</p><p>This separation of concerns, that is, <strong>Elasticsearch handling data integration and retrieval, and agents focusing on reasoning, planning, and decision-making</strong>,is what makes agent systems scalable, reliable, and production ready.</p><h2><strong>Elastic’s role in the AI stack</strong></h2><p>Elastic sits at the intersection of search and AI by design.</p><ul><li><p><strong>Connectors and crawlers</strong> ingest data continuously from enterprise systems.</p></li><li><p><strong>Semantic and vector search</strong> enable intent-based retrieval.</p></li><li><p><strong>Hybrid search</strong> blends lexical precision with semantic understanding.</p></li><li><p><strong>RAG workflows</strong> ground LLMs in authoritative, permission-aware data.</p></li></ul><p>Elastic does not compete with agents or MCP. It <strong>makes them effective</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt950398e2b4ab71f8/6a170cac2b835fb826f4b26b/193da239544ce858416db845f9fc34c7c0e9b6f9-1920x1080.png" alt="AI-native experiences powered by tools and agents, enabled by the platform, and built on enterprise data." /><h2><strong>Bigger models don’t eliminate retrieval</strong></h2><p>Some have wondered whether huge new LLMs can bypass traditional search, perhaps by letting the model read <em>everything</em> in one go. Large context windows feel powerful, but they introduce:</p><ul><li><p>Higher latency.</p></li><li><p>Higher cost.</p></li><li><p>Lower precision due to noise.</p></li><li><p>A higher propensity for confusion, context clash, and context poisoning.</p></li></ul><p>RAG wins because it filters first and then reasons.In another <a href="https://www.elastic.co/search-labs/blog/rag-vs-long-context-model-llm#:~:text=,context%20approach%20led%20to%20inaccuracies">Elastic Search Labs experiment</a>, RAG achieved answers in about <strong>1 second</strong>, versus 45 seconds for the raw-LM approach, at <strong>1/1250th</strong> the cost, and with far higher accuracy. In other words, giving an LLM a million tokens of documents is slower, more expensive, and actually <em>less precise</em> than filtering through an index first.</p><h2><strong>Conclusion: MCP changes the interface, not the fundamentals</strong></h2><p>MCP is a meaningful step forward in how agents interact with tools. But it doesn’t replace the need for fast, relevant, governed retrieval.</p><p>In enterprise AI:</p><ul><li><p>Context quality determines answer quality.</p></li><li><p>Indexes create that context.</p></li><li><p>Search is the foundation, not the legacy.</p></li></ul><p>Indexes aren’t obsolete in the era of MCP. They’re <strong>the reason that MCP-based agents can work at all</strong>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/future-of-search-engines-indexed-search-mcp</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/future-of-search-engines-indexed-search-mcp</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Relevance]]></category>
    <category><![CDATA[Agentic AI]]></category>
    <dc:creator><![CDATA[Dayananda Srinivas]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1caa3ee906789415/6a170cae2b835f9a90f4b26f/5b8af1c3ca51f2c038406c714eb9a71b696bbc5a-1999x1091.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 05 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Adaptive early termination for HNSW in Elasticsearch]]></title>
    <description><![CDATA[Introducing a new adaptive early termination strategy for HNSW in Elasticsearch.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch uses the <a href="https://www.elastic.co/search-labs/blog/hnsw-graph">Hierarchical Navigable Small World</a> (HNSW) algorithm to perform vector search over a proximity graph. HNSW is known to provide a nice trade-off between the quality of k-nearest neighbor (KNN) results and the associated cost.</p><p>In HNSW, search proceeds by iteratively expanding candidate nodes in the graph, maintaining a bounded set of nearest neighbors discovered so far. Each expansion has a cost (vector operations, random seeks to disk, and more), and the marginal benefit of that cost tends to decrease as the search progresses.</p><p>One way to optimize HNSW graph traversal is to stop searching when the marginal likelihood of finding new true neighbors doesn’t increase. For this reason, in <a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/index-modules#index-dense-vector-hnsw-early-termination">Elasticsearch 9.2</a> we introduced a new <a href="https://www.elastic.co/search-labs/blog/hnsw-knn-search-early-termination">early termination mechanism</a>. This stops the search process when visiting graph nodes doesn’t provide enough new nearest neighbors, consecutively, for a fixed number of times.</p><p>This article guides you through how we improved over the mentioned early termination mechanism in HNSW to make it better suited for different datasets and data distributions.</p><h2><strong>Early termination in HNSW</strong></h2><p>In HNSW, search proceeds by iteratively expanding candidate nodes in the proximity graph, maintaining a bounded set of nearest neighbors discovered so far, until it either has visited the whole graph or meets some early stop criteria.</p><p>Early termination is therefore not necessarily always an optimization, it’s <strong>part of the search algorithm itself</strong>. The moment we decide to stop determines the balance between efficiency and recall. In Elasticsearch, there are already a number of ways a query on HNSW can early terminate:</p><ul><li><p>A fixed maximum number of nodes is visited.</p></li><li><p>A fixed timeout is reached.</p></li></ul><p>While simple and predictable, these rules are largely <strong>agnostic to what the search is actually doing</strong>. Also they’re used mostly to make sure that the query finishes in reasonable time for the end user.</p><p>In a <a href="https://www.elastic.co/search-labs/blog/hnsw-knn-search-early-termination">previous blogpost</a>, we introduced the concept of redundancy in HNSW. In short, redundant computations occur when HNSW continues to evaluate new candidate nodes that don’t result in finding more nearest neighbors.</p><h2><strong>Patience: Measuring progress instead of effort</strong></h2><p>The notion of <em>patience</em> reframes early termination around <strong>progress rather than effort</strong>.</p><p>Instead of asking:</p><p>“How many steps have we taken?”</p><p>The new question becomes:</p><p>“What is the amount of computation we accept to waste, until we lose hope?”</p><p>During HNSW search, early exploration typically produces peak improvements to the top-k candidate set. During first steps of the HNSW graph exploration, the set of neighbors is continuously updated as the algorithm keeps discovering nearer and nearer neighbors to the query vector. Over time, these improvements become rarer as the search converges. <a href="https://cs.uwaterloo.ca/~jimmylin/publications/Teofili_Lin_ECIR2025.pdf">Patience-based termination</a> monitors this pattern and terminates the search once improvements have ceased for a sustained period.</p><p>In practice, while visiting the HNSW graph we also compute the queue saturation ratio as we hop through candidate nodes. This measures the percentage of nearest neighbors that were left unchanged while visiting the most recent graph node (or the inverse of the number of new neighbors introduced during the last iteration). When such a ratio becomes too big for too many consecutive iterations, we stop visiting the graph.</p><p>Conceptually, patience treats HNSW search as a <strong>diminishing returns process</strong>. When returns flatten out, continuing to explore the graph yields little benefit.</p><p>This framing is powerful because it ties termination directly to <em>observable outcomes</em> rather than to arbitrary fixed limits.</p><p>The benefit of using this smart early termination technique is that HNSW graph explorations tend to visit a smaller number of graph nodes while retaining an almost perfect relative recall.</p><p>To visualize this, we can plot the amount of recall per visited node that we got with the patience based early termination (labeled as <em><code>et=static</code></em>), when compared to the default HNSW behavior (labeled as <em><code>et=no</code></em>) on a couple of datasets, FinancialQA and Quora, and models, JinaV3 and E5-small.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfd0d692b9beb476a/6a170ef4dc55debf0be00e97/a9d07c5153ea64a2426c82487c36846030692bb9-1600x945.png" alt="Adaptive Early Termination for HNSW " /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt93509c251a1b641e/6a170ef6dc55dea2b3e00e9b/dac56125c4b16d1b596c9876b6ca9ac7b2dc87fa-1600x944.png" alt="Adaptive Early Termination for HNSW es" /><h2><strong>Static thresholds and HNSW dynamics</strong></h2><p>In practice, in Elasticsearch this is implemented using <strong>static thresholds</strong>. One threshold refers to the <strong>saturation threshold</strong>: that is, the ratio of saturation that we consider suboptimal. The other threshold refers to the number of consecutive graph nodes that we allow to be visited while still having a suboptimal queue saturation: that is, the <strong>patience threshold</strong>.</p><p>When we introduced this early termination strategy in Elasticsearch 9.2, we decided to opt for conservative defaults, so as to let the recall as much as possible, while still gaining in terms of latency and memory consumption. For this reason, we set the saturation threshold to be 100% and the patience threshold to be set as a (bounded) 30% of the <a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-knn-query#knn-query-top-level-parameters:~:text=search%20request%20size.-,num_candidates,-(Optional%2C%20integer)%20The"><em><code>num_candidates</code></em></a> in the KNN query.</p><p>In many scenarios, these settings resulted to work nicely; however, two queries requesting the same number of neighbors might have radically different convergence behaviors. Some queries encounter dense local neighborhoods and saturate quickly; others must traverse long, sparse paths before finding competitive candidates. The latter resulted to be the most difficult to handle effectively.</p><p>As a result, we sometimes noticed:</p><ul><li><p>Over-exploration for easy queries.</p></li><li><p>Premature termination for hard queries.</p></li></ul><p>Therefore, we figured that fixed threshold values encode global assumptions about convergence, whereas we could make HNSW better adapt to different dynamics.</p><h2><strong>Making HNSW early termination adaptive</strong></h2><p>Adaptive early termination approaches this problem from a different angle. Instead of enforcing predefined stopping thresholds, the algorithm <strong>infers when to stop from the search dynamics themselves</strong>.</p><p>So instead of comparing the queue saturation ratio between two consecutive candidates, we decided to introduce both an instant smoothed discovery rate   (how many new neighbors were introduced for a query <em>q</em>, in the last visit <em>i</em>) together with rolling mean  and standard deviation  of such a discovery rate during the graph visit (using <a href="https://en.wikipedia.org/wiki/Algorithms_for_calculating_variance#Welford's_online_algorithm">Welford’s algorithm</a>). These statistics about the discovery rate are calculated per query, so that this information can be used to decide different degrees of patience for each query.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdbfb1e123f1b026d/6a170ef7cf4f25d9bab2d216/1958be7ca4425ade66eaf621ada3533173183598-694x118.png" alt="" /><p>The previously static thresholds become adaptive to the discovery rate statistics: The saturation threshold becomes the rolling mean plus the standard deviation; whereas we make the patience adapt and scale inversely with the standard deviation.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7d4d91f464a9dc9d/6a170ef8d7c0223420de656a/f7ee4a55c24853b657df26052b275e8bd76cf0f9-654x156.png" alt="" /><p>The early exit rules remain the same; the saturation happens when the instant discovery rate is lower than the adaptive saturation threshold. The graph visit stops if the saturation persists for a number of consecutive candidate visits that’s larger than the adaptive patience.</p><p>This way, we obtain a behavior that doesn’t depend on the <em><code>num_candidates</code></em> parameter in the KNN query (which might be always set or left as the default, regardless of early exit) and that better adapts to each query and vector distribution dynamically.</p><p>The recall per visited node on FinancialQA and Quora with the adaptive strategy (labeled as <em><code>et=adaptive</code></em>) reports a higher recall per visited node, when compared to the static strategy (<em><code>et=static</code></em>) and the default HNSW behavior (<em><code>et=no</code></em>).</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blteab7ba53ae14da0e/6a170ef9961e69e072c4cfd5/2a906997d9a25d74c7038bd9661bc97581e7258e-1600x938.png" alt=" adaptive strategy and the default HNSW behavior" /><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6fb9e672d3698200/6a170efb67045b7b2b45c2ab/3a114911e232c351dbb814cea20e8b0f1415a717-1600x925.png" alt="" /><p>Adaptive early termination is turned on by default in Elasticsearch 9.3 for HNSW dense vector fields (and it can eventually be turned off via the <a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/index-modules#index-dense-vector-hnsw-early-termination">same index level setting</a>).</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/hnsw-elasticsearch-adaptive-early-termination</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/hnsw-elasticsearch-adaptive-early-termination</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Tommaso Teofili]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt27b746cc1995e6b7/6a170efda29299de8ad010c6/e6d3186f609dd56dc5ffe33d70fa9e5cfa05b51f-1280x720.png" length="0" type="image/png"/>
    <pubDate>Mon, 02 Mar 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Speed up vector ingestion using Base64-encoded strings]]></title>
    <description><![CDATA[Introducing Base64-encoded strings to speed up vector ingestion in Elasticsearch.]]></description>
    <content:encoded><![CDATA[<p>We’re improving the ingestion speed of vectors in Elasticsearch. Now, in <a href="https://www.elastic.co/cloud/serverless">Elastic Cloud Serverless</a> and in v9.3, you can send your vectors to Elasticsearch encoded as Base64 strings, which will provide immediate benefits to your ingestion pipeline.</p><p>This change reduces the overhead of parsing vectors in JSON by an order of magnitude, which translates to almost a 100% improvement on indexing throughput for DiskBBQ and around 20% improvement for hierarchical navigable small world (HNSW) workloads. In this blog, we’ll take a closer look at Base64-encoded strings and the improvements it brings to vector ingestion.</p><h2>What’s the problem?</h2><p>At Elastic, we’re always looking for ways to improve our vector search capabilities, whether that’s enhancing existing storage formats or introducing new ones. Recently, for example, we added a new disk-friendly storage format called <a href="https://www.elastic.co/search-labs/blog/diskbbq-elasticsearch-introduction">DiskBBQ</a> and enabled vector indexing with <a href="https://www.elastic.co/search-labs/blog/elasticsearch-gpu-accelerated-vector-indexing-nvidia">NVIDIA cuVS</a>.</p><p>In both cases, we expected to see major gains in ingestion speed. However, once these changes were fully integrated into Elasticsearch, the improvements weren’t as large as we had hoped. A flamegraph of the ingestion process made the issue clear: JSON parsing had become one of the main bottlenecks.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7d4cd9d8628b64b3/6a170c5b2867148d5e93e353/a286408afc85ff1cd3dd448b8fdf59dd3e11d599-1600x675.png" alt="Vector ingestion before using Base64-encoded strings  " /><p>Parsing JSON requires walking through every element in the arrays and converting numbers from text format into 32-bit floating-point values, which is very expensive.</p><h3>Why Base64-encoded strings?</h3><p>The most efficient way to parse vectors is directly from their binary representation, where each element uses a 32-bit floating-point value. However, JSON is a text-based format, and the way to include binary data in it is by using <a href="https://en.wikipedia.org/wiki/Base64">Base64</a>-encoded strings. Base64 is just a binary-to-text encoding schema.</p>{
  “emb” : [1.2345678, 2.3456789, 3.4567891]
}<p>We can now send vectors encoded as Base64 strings:</p>{
  “emb” : ”P54GUUAWH5pAXTwI”
}<p>Is it worth it? Our benchmarks suggest yes. When parsing 1,000 JSON documents, using Base64 encoded strings instead of float arrays resulted in performance improvements of more than an order of magnitude, at the cost of a small encode/decode trade-off (client-side Base64 encoding and a temporary byte array on the server for decoding) in exchange for eliminating expensive per-element numeric parsing.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta9a1662fdf7d5849/6a170c5d839dfaf624dcff29/86e5a926e13b07bb3b0abe80bd4930464e8f6f9b-1200x742.png" alt="Base64 vs. Float32 parsing time" /><h3>Give me some ingestion numbers</h3><p>We can see these improvements in practice when running the <a href="https://github.com/elastic/rally-tracks/blob/master/so_vector/README.md"><code>so_vector</code></a> rally track with the different approaches. The actual gains depend on how fast indexing is for each storage format. For <code>bbq_disk</code>, indexing throughput increases by about 100%, while for <code>bbq_hnsw</code>, the improvement is closer to 20%, since indexing is inherently slower there.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte35ffc920ae863f5/6a170c5e509168f193e1bb1c/4277057ee59cb84d068176b56bb7fa00b66e1cb3-1200x742.png" alt="Base64 vs Float32 indexing throughput" /><p>Starting with Elasticsearch v9.2, <a href="https://www.elastic.co/search-labs/blog/elasticsearch-exclude-vectors-from-source">vectors are excluded from </a><a href="https://www.elastic.co/search-labs/blog/elasticsearch-exclude-vectors-from-source"><code>_source</code></a> by default and are stored internally as 32-bit floating-point values. This behavior also applies to Base64-encoded vectors, making the choice of indexing format completely transparent at search time.</p><h2>Client support</h2><p>Adding a new format for indexing vectors might require changes on ingestion pipelines. To help this effort, in v9.3, Elasticsearch official clients can transform vectors with 32-bit floating-point values into Base64-encoded strings and the other way around. You might need to check the client documentation for the specific implementation.</p><p>For example, here’s a snippet for implementing bulk loading using the Python client:</p>from elasticsearch.helpers import bulk, pack_dense_vector

def get_next_document():
    for doc in dataset:
        yield {
            "_index": "my-index",
            "_source": {
                "title": doc["title"],
                "text": doc["text"],
                "emb": pack_dense_vector(doc["emb"]),
            },
        }

result = bulk(
    client=client,
    chunk_size=chunk_size,
    actions=get_next_document,
    stats_only=True,
)<p>The only difference from a bulk ingest using floats is that the embedding is wrapped with the <code>pack_dense_vector()</code> auxiliary function.</p><h2>Conclusion</h2><p>By switching from JSON float arrays to Base64-encoded vectors, we remove one of the largest remaining bottlenecks in Elasticsearch’s vector ingestion pipeline: numeric parsing. The result is a simple change with outsized impact: up to 2× higher throughput for DiskBBQ workloads and meaningful gains even for slower indexing strategies, like HNSW.</p><p>Because vectors are already stored internally in a binary format and excluded from <code>_source</code> by default, this improvement is completely transparent at search time. With official client support landing in v9.3, adopting Base64 encoding requires only minimal changes to existing ingestion code, while delivering immediate performance benefits.</p><p>If you’re indexing large volumes of embeddings, especially in high-throughput or serverless environments, Base64-encoded vectors are now the fastest and most efficient way to get your data into Elasticsearch.Those interested in the implementation details can follow the related Elasticsearch issues and pull requests: #<a href="https://github.com/elastic/elasticsearch/issues/111281">111281</a> and #<a href="https://github.com/elastic/elasticsearch/issues/135943">135943</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/base64-encoded-strings-vector-ingestion</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/base64-encoded-strings-vector-ingestion</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Jim Ferenczi,Benjamin Trent,Ignacio Vera Sequeiros]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc5ffc7ac4c2b9d93/6a170c5f839dfa007ddcff2d/4c1ebbd7a1071e8e1721a9871cba87f6aed140e9-1280x720.png" length="0" type="image/png"/>
    <pubDate>Wed, 04 Feb 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Influencing BM25 ranking with multiplicative boosting in Elasticsearch]]></title>
    <description><![CDATA[Learn why additive boosting methods can destabilize BM25 rankings and how multiplicative scoring provides controlled, scalable ranking influence in Elasticsearch.]]></description>
    <content:encoded><![CDATA[<p><a href="https://en.wikipedia.org/wiki/Okapi_BM25">BM25</a> is one of the most widely used scoring models in Elasticsearch for text-based search. In many e-commerce implementations, it forms a major component of how product relevance is determined because it provides a well-understood, interpretable score that reflects how closely an item matches a shopper’s query. In addition to this text relevance, merchandising and search teams often need to influence the ranking with business metrics such as margin, stock levels, popularity, personalization, or campaign strategy, in a way that doesn’t destabilize the underlying text relevance.</p><p>The most intuitive levers for doing this are boosted <a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-bool-query">should</a> clauses or <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/rank-feature">rank_feature</a> fields. These may initially appear effective, but both approaches degrade and may even fail, as query patterns shift or catalog composition changes. Their shared limitation is that they introduce additive adjustments into a scoring system whose scale varies substantially across queries. A boost like “+2” might overwhelm the base BM25 score in one query while barely registering in another. In other words, additive methods may create brittle, unpredictable ranking behavior.</p><p>In contrast, <a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-function-score-query">function_score</a> with multiplicative boosting provides a stable and mathematically proportional way to shape BM25 scores without distorting their underlying structure. Your application logic determines what merits uplift; <code>function_score</code> expresses that intent in a predictable and explainable way that preserves the geometry (high-level relative ordering) of the BM25 relevance signal, nudging rankings in controlled ways rather than overwhelming the core text relevance.</p><p>This article builds on two earlier pieces that demonstrated practical uses of multiplicative boosting: (1) <a href="https://www.elastic.co/search-labs/blog/function-score-query-boosting-profit-popularity-elasticsearch">Boosting e-commerce search by profit and popularity with the function score query in Elasticsearch</a>, and (2) <a href="https://www.elastic.co/search-labs/blog/ecommerce-search-relevance-cohort-aware-ranking-elasticsearch">How to improve e-commerce search relevance with personalized cohort-aware ranking</a>. Here we step back from those examples to examine the architectural principle that underlies them: why multiplicative boosting via <code>function_score</code> is one of the most reliable and scalable ways to influence BM25-based ranking in Elasticsearch.</p><h2>Why it's important to preserve base BM25 rankings</h2><p>In many Elasticsearch-based applications, including e-commerce, BM25 remains a central component of how text relevance is assessed. It provides a signal that is interpretable and transparent for teams who need to understand why a product ranked where it did. These properties make BM25 particularly attractive in environments where explainability and operational predictability matter.</p><p>Because of this, most teams want to shape, rather than replace, the rankings produced by BM25. For example, they may want to allow higher-margin items to surface slightly more often, reduce exposure for low-stock products without hiding them, or highlight items aligned with a particular user segment. Ideally, this shaping should preserve the geometry of the rankings produced by the BM25 algorithm.</p><p>The difficulty arises when teams try to achieve these goals using mechanisms that add separate scoring streams on top of the base BM25 ranking. These additive adjustments are not always comparable to BM25’s scale and behave inconsistently as queries, data distributions, and catalog composition evolve. Over time, the ranking becomes brittle, unintuitive, and difficult to tune. A reliable influence mechanism must work with BM25’s scoring geometry rather than overpowering it.</p><p>The <a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-function-score-query">function_score</a> query with multiplicative boosting provides this property. It allows teams to apply business influence in a proportional, explainable way while keeping BM25’s underlying structure intact.</p><h2>Why many approaches to influencing ranking degrade (or break) BM25</h2><p>Teams often begin with mechanisms that look straightforward: boosted <a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-bool-query">should</a> clauses, <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/rank-feature">rank_feature</a> fields, or custom <a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-script-score-query">script_score</a> logic. These tools can be effective in their intended use cases, which is why they seem like natural levers for adding business influence. But when they are used to shape or influence BM25-based text relevance, they may create unstable, opaque, or brittle ranking behavior.</p><p>The underlying issue is that these approaches introduce independent additive scoring contributions into a system whose base BM25 values vary widely across queries, fields, and data sets. Without respecting that variability, the influence becomes unpredictable.</p><p>Below are the three most common patterns and why they fail in practice.</p><h3>1. Additive boosts via should clauses</h3><p>A boosted <code>should</code> clause feels intuitive: “Promote items that match this business rule.” But under the hood, the behavior is fundamentally additive.</p><p>Consider a query of the form:</p>GET products/_search
{
  "query": {
    "bool": {
      "must": [ { "match": { "description": "running shoes" }}],
      "should": [ { "term": { "brand": { "value": "nike", "boost": 1 }}}]
    }
  }
}<p>This kind of query results in the following behavior:</p>final_score = base_BM25 + should_BM25<p>The problem is that <code>base_BM25</code> and <code>should_BM25</code> do not scale together. As your dataset changes, or as different queries are issued, the magnitude of BM25 can shift dramatically. For example, the base BM25 scores for three products might be 12, 8, 4 in one context, and 0.12, 0.08, 0.04 in another. Such a change might happen after a catalog update or a modification to the query structure.</p><p>A boosted <code>should</code> clause adds its own BM25-style contribution to the final score. In this situation, an additive contribution (i.e. should_BM25 = +2) behaves inconsistently:</p><ul><li><p>When base_BM25 is small (0.12), +2 dominates the score — roughly an 18× increase.</p></li><li><p>When base_BM25 is large (12), the same +2 barely shifts the document —  only about a 17% increase.</p></li></ul><p>This instability means that the combined <code>must</code> score and <code>should</code> score have no stable meaning across queries or catalogs. A rule that slightly promotes a brand for one query can dominate the ranking for another, or become irrelevant in a third. This is not a tuning issue; it is a structural property of additive scoring.</p><h3>2. Using rank_feature for business influence</h3><p>The <code>rank_feature</code> family is extremely useful for representing numeric qualities such as recency or popularity. It is fast, compressed, and operationally simple. However, when it is used to influence text relevance (BM25), it runs into the same structural limitation described in the previous section.</p><p>A <code>rank_feature</code> clause produces its own scoring contribution, which is then added to the BM25 score:</p>final_score = base_BM25 + feature_score<p>Just as with boosted <code>should</code> clauses, the two components do not scale together. BM25 values vary substantially across queries depending on term rarity and catalog statistics, while the <code>feature_score</code> follows the scale of the underlying business attribute being boosted (for example, popularity or recency), which typically bears no relationship to the scale of BM25. As a result, the two scoring streams drift apart as your corpus or query patterns evolve.</p><p>The consequence is the same as what we discussed above with relation to the should-clause problem:</p><ul><li><p>The feature score can dominate BM25 in one query and be negligible in another.</p></li><li><p>Tuning becomes fragile because you are calibrating two independent scales — BM25, which varies with query term statistics, and the feature score, which varies with the business attribute’s own distribution.</p></li></ul><p>Although <code>rank_feature</code> remains an excellent mechanism for representing raw numeric attributes, it is not well-suited for proportional influence on BM25, where the goal is not to add a second score but to gently shape the existing one.</p><h3>Custom scoring with script_score</h3><p>When boosted clauses or <code>rank_feature</code> fields become difficult to tune, teams often turn to <code>script_score</code> as a last resort. It provides complete freedom to manipulate the score, including adding, subtracting, multiplying, or replacing the BM25 value according to any business rule. A <code>script_score</code> query replaces Elasticsearch’s scoring pipeline with custom logic. Instead of shaping the BM25 score, the script builds a separate scoring mechanism whose behavior depends entirely on the code inside the script. While this can be powerful, it introduces three challenges that become more significant as the system grows.</p><p><strong>1. Opacity</strong></p><p>Scoring logic is hidden inside a script rather than expressed declaratively. When ranking behavior changes unexpectedly, it is difficult to understand whether the issue is the script itself, a data shift, or an interaction with BM25. Merchandisers and relevance engineers lose the ability to reason about why a document moved up or down.</p><p><strong>2. Performance and operational cost</strong></p><p>Script scoring bypasses many of Elasticsearch’s optimizations and caching pathways. Each document that matches the initial query must execute the script, often leading to higher CPU usage and unpredictable latency.</p><p><strong>3. Fragility when combined with BM25</strong></p><p>Because <code>script_score</code> allows arbitrary computations, it is easy to drift into scoring behaviors that no longer resemble BM25 or that fail to preserve its relative structure. As the dataset evolves or query patterns shift, the custom logic may interact with BM25 in unanticipated ways. A script that behaved reasonably early in development can produce surprising or unstable results once the catalog grows or data distributions change. Because <code>script_score</code> allows arbitrary math, two engineers working on different parts of the system may unintentionally encode competing scoring models, making ranking difficult to reason about as the organization scales.</p><h2>How function_score provides predictable influence on BM25</h2><p>BM25 already captures how well a document matches a query. It reflects text relevance, term rarity, document length, and the statistical shape of the corpus. When teams introduce business signals including margin, stock levels, popularity, personalization, or merchandising strategy, the goal is not to replace this relevance. The goal is to <em>influence it</em>.</p><p>This distinction is subtle but crucial. Most business requirements are proportional in nature:</p><ul><li><p>Promote higher-margin items modestly</p></li><li><p>Reduce exposure for low-stock products, but don’t hide them</p></li><li><p>Give this user segment a slight uplift for matching products</p></li><li><p>Boost for popularity, but not so much that textual relevance is lost</p></li></ul><p>These are naturally expressed as <em>percentage adjustments</em> rather than as fixed additive values. A merchandiser is rarely asking for “+2 points of score”; they are asking for “a little more visibility,” irrespective of the absolute numeric scale of the BM25 score. Mathematically, this means that the desired transformation is:</p>final_score = BM25 × boost_factor<p>Where <em>boost_factor</em> might be 1.05, 1.2, or 1.5, depending on the signal. Multiplicative boosting does not attempt to reinvent scoring; it simply adjusts the BM25 output by a proportional factor. A multiplicative adjustment has three properties that align well with real-world ranking control:</p><ol><li><p>The boost remains proportional. In other words, a 20% uplift is always a 20% uplift—whether BM25 is 0.12 or 12. The magnitude of the boost does not depend on the underlying BM25 scale.</p></li><li><p>BM25 retains its role as the primary signal. The multiplicative shaping nudges the ordering without overriding it. Strong textual matches still win; business logic influences but does not dominate.</p></li><li><p>Because the operation is multiplicative, not additive, changing the query or updating the corpus does not require re-tuning numeric constants. The boost has the same meaning everywhere.</p></li></ol><p>Elasticsearch’s <code>function_score</code> query provides an elegant mechanism for expressing this pattern. By using:</p><ul><li><p><strong>score_mode: “sum”</strong> to assemble a boost factor (building the multiplier), and</p></li><li><p><strong>boost_mode: “multiply”</strong> to apply the boost (multiplier) to BM25</p></li></ul><p>You can express business intent in a way that remains stable and explainable as your data and query patterns evolve. Instead of adding a second score beside BM25, <code>function_score</code> transforms BM25 itself—shaping it gently, predictably, and in line with how merchandisers and product owners think about ranking adjustments.</p><h2>Examples in practice: How multiplicative boosting behaves in real e-commerce queries</h2><p>To illustrate how multiplicative boosting works in real-world ranking scenarios, it helps to look at a small, concrete example. The goal here is not to demonstrate tuning or production-scale scoring, but rather to show how <code>function_score</code> influences BM25 in predictable, proportional ways that align with business intent.</p><p>Consider a simple catalog with three basketball shoes from three different brands: Nike, Adidas, and Reebok. The product descriptions are intentionally crafted so the BM25 scores exhibit natural differences based on query specificity and field length—just as they would in a real catalog.</p><h3>Example dataset</h3><p>For the following examples, we use a small, straightforward sample dataset with the following characteristics.</p><p>Brand</p><p>Description</p><p>nike</p><p>“Nike basketball shoes”</p><p>adidas</p><p>“New Adidas basketball shoes”</p><p>reebok</p><p>“Reebok basketball shoes”</p><p>We can create an index with the above products with the following commands from Kibana Dev Tools:</p>PUT products
{
  "mappings": {
    "properties": {
      "brand":       { "type": "keyword" },
      "description": { "type": "text" }
    }
  }
}

POST products/_bulk
{ "index": { "_id": "nike-001" } }
{ "brand": "nike",    "description": "Nike basketball shoes" }
{ "index": { "_id": "adi-001" } }
{ "brand": "adidas",  "description": "New Adidas basketball shoes" }
{ "index": { "_id": "ree-001" } }
{ "brand": "reebok", "description": "Reebok basketball shoes" }<p>With this dataset, we now evaluate three queries:</p><ul><li><p>A baseline “basketball shoes” search</p></li><li><p>The same query with a 50% promotion for Adidas and a 25% promotion for Nike</p></li><li><p>A specific “Reebok basketball shoes” query while the Adidas and Nike promotions are still active</p></li></ul><p>Each scenario highlights a different property of multiplicative boosting.</p><h3>1. Baseline ranking: No promotion</h3>GET products/_search
{
  "size": 3,
  "_source": ["brand", "description"],
  "query": {
    "match": { "description": "basketball shoes" }
  }
}<img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt272891c9a78166c1/6a16f9d12b835f9e90f4b00b/050f44956c112ea9916d34946e9296480354f7d3-2322x1324.png" alt="" /><p>This query returns the following results where Nike and Reebok are ranked above adidas:</p><p>Rank</p><p>Brand</p><p>Score (BM25)</p><p>1/2 (tie)</p><p>nike</p><p>0.27845407</p><p>1/2 (tie)</p><p>reebok</p><p>0.27845407</p><p>3</p><p>adidas</p><p>0.24686474</p><h3>2. Adding 50% Adidas uplift and 25% Nike uplift with function_score</h3><p>If marketing launches a campaign where Adidas basketball shoes should receive a 50% uplift and Nike a 25% uplift, then the application layer could construct its queries to include those uplifts as follows:</p>GET products/_search
{
  "size": 3,
  "_source": ["brand", "description"],
  "query": {
    "function_score": {
      "query": {
        "match": { "description": "basketball shoes" }
      },
      "functions": [
        {
          "filter": { "term": { "brand": "adidas" } },
          "weight": 0.5
        },
        {
          "filter": { "term": { "brand": "nike" } },
          "weight": 0.25
        },
        {
          "weight": 1.0
        }
      ],
      "score_mode": "sum",
      "boost_mode": "multiply"
    }
  }
}<h3>How the multiplier is constructed</h3><ul><li><p>Base weight = 1.0</p></li><li><p>Adidas gets an additional +0.5</p></li><li><p>So Adidas’s multiplier = 1.5</p></li><li><p>Nike gets an additional +0.25</p></li><li><p>So Nike’s multiplier = 1.25</p></li><li><p>All other brands (including Reebok) get the base weight multiplier = 1.0</p></li></ul><h3>Apply multiplier:</h3><p>Final score = BM25 × multiplier</p><p>Product</p><p>BM25</p><p>Multiplier</p><p>Final score</p><p>Adidas</p><p>0.24686474</p><p>1.5</p><p>0.37029710</p><p>Nike</p><p>0.27845407</p><p>1.25</p><p>0.34806758</p><p>Reebok</p><p>0.27845407</p><p>1.0</p><p>0.27845407</p><h3>Result</h3><p>Adidas moves to the top, Nike follows, and Reebok is at the bottom with no change in its score. This is exactly the behavior that multiplicative boosting is designed to produce:</p><ul><li><p>Adidas and Nike both gain visibility, but in proportion to their configured uplifts.</p></li><li><p>The relative differences in BM25 still matter; we are reshaping the ranking, not replacing it.</p></li><li><p>The ordering changes primarily where BM25 scores are close.</p></li></ul><p>With additive boosts, the same “50% versus 25%” business intent would have to be approximated with numeric constants on an arbitrary BM25 scale, and the effect would vary drastically across queries.</p><h2>3. Specific intent still wins: “Reebok basketball shoes”</h2><p>Now run a highly specific branded query for “Reebok basketball shoes”, with the same Adidas (50%) and Nike (25%) promotions still active:</p>GET products/_search
{
  "size": 3,
  "_source": ["brand", "description"],
  "query": {
    "function_score": {
      "query": {
        "match": { "description": "Reebok basketball shoes" }
      },
      "functions": [
        {
          "filter": { "term": { "brand": "adidas" } },
          "weight": 0.5
        },
        {
          "filter": { "term": { "brand": "nike" } },
          "weight": 0.25
        },
        {
          "weight": 1.0
        }
      ],
      "score_mode": "sum",
      "boost_mode": "multiply"
    }
  }
}<p>The response shows the following results:</p><p>Rank</p><p>Brand</p><p>Final score</p><p>1</p><p>reebok</p><p>1.3011196</p><p>2</p><p>adidas</p><p>0.3702971</p><p>3</p><p>nike</p><p>0.34806758</p><h3>Result</h3><p>Reebok wins overwhelmingly because BM25 correctly detects strong intent for “Reebok basketball shoes”. Adidas and Nike still receive their 50% and 25% promotions, respectively, but those multipliers are nowhere near enough to override the BM25 score.</p><p>This is exactly the behavior that multiplicative boosting is designed to produce:</p><ul><li><p>When BM25 scores are close, boosts can shift the relative ordering.</p></li><li><p>When BM25 scores differ significantly (as they do here, due to strong text matching), the same boosts have little practical effect.</p></li></ul><p>Promotions influence the ranking, but they do not override the core text relevance signal.</p><h2>What this example demonstrates</h2><p>These real queries illustrate the key properties of multiplicative boosting:</p><ol><li><p>The influence is proportional, not arbitrary. A percentage-based uplift has the same proportional effect regardless of the underlying BM25 scale.</p></li><li><p>Text relevance remains in control. Strong brand-intent queries still surface the correct product.The system behaves intuitively. Merchandisers see exactly the ranking changes they expect.</p></li><li><p>The math is stable across queries. The same promotion works correctly whether the match is broad or highly specific.</p></li><li><p>Application logic stays clean. The business layer decides the uplift; Elasticsearch applies it predictably.</p></li></ol><p>Multiplicative boosting through <code>function_score</code> preserves relevance in a predictable and controllable way, while enabling business impact.</p><h2>Application logic remains the author of influence</h2><p>There is a clear separation between deciding what should be boosted and applying that boost in Elasticsearch. <code>function_score</code> handles the second task, but the first belongs firmly to application logic.</p><p>Your application logic is where decisions are made about:</p><ul><li><p>Which margin thresholds matter for your business</p></li><li><p>Whether popularity should rise or fall based on seasonality</p></li><li><p>How to interpret customer behavior or cohort membership</p></li><li><p>How to encode campaign rules</p></li><li><p>When to surface or suppress certain product groups</p></li></ul><p>These are <em>business</em> decisions, not scoring decisions. Elasticsearch does not infer whether a user is budget-focused or luxury-oriented, whether a promotion is active, or whether low stock requires a visibility adjustment. Those determinations occur upstream, in the part of the system that has access to user context, session features, analytics, and business configuration. After application logic produces clear numeric signals for fields such as weights, uplift factors, thresholds, and cohort tags, a <code>function_score</code> query provides a reliable way to express those signals as controlled multipliers on BM25.</p><p>This creates a clean architectural contract:</p><ul><li><p>Application logic: decides <em>what</em> should be influenced.</p></li><li><p>BM25 provides the core text relevance.</p></li><li><p><code>function_score</code> applies influence in a mathematically stable way.</p></li></ul><p>Because business logic lives outside the index, teams can adjust or experiment with uplift strategies without reindexing or restructuring documents.</p><h2>Conclusion</h2><p>E-commerce search must balance core text relevance with business considerations such as profitability, stock position, customer intent, seasonality, and personalization. BM25 provides a stable and interpretable foundation for text relevance, but influencing that score requires care. Business signals should shape the ranking, not overpower it.</p><p>However, the most commonly used levers such as boosted <code>should</code> clauses, <code>rank_feature</code> fields, and ad-hoc script scoring often behave unpredictably. These approaches can appear effective in early development, but their limitations emerge as soon as the catalog evolves or new query patterns arrive. Additive boosts fluctuate wildly because their impact depends entirely on the underlying scale of BM25, which varies dramatically across queries. A boost that produces a subtle nudge in one situation can dominate the ordering in another. Script scoring introduces its own challenges: opaque logic, reduced performance, and scoring behavior that becomes harder to understand or maintain over time.</p><p>Multiplicative boosting with <code>function_score</code> avoids these pitfalls by transforming BM25 proportionally rather than competing with it. Instead of adding a second, independent score component, it applies a controlled multiplier to BM25 itself. This produces the kind of predictable adjustments that merchandisers actually intend. For example, it allows slight promotions for high-margin items, modest reductions for low-stock products, or gentle uplifts for relevant user cohorts.</p><p>Equally important, the architecture remains clean. Application logic determines which business signals matter, and <code>function_score</code> applies them in a consistent, explainable way. Business teams can evolve business strategy without destabilizing relevance, and Engineering teams can refine relevance without disturbing business rules.</p><p>This principle is the foundation of the previous blogs that demonstrated how to influence e-commerce rankings: (1) <a href="https://www.elastic.co/search-labs/blog/function-score-query-boosting-profit-popularity-elasticsearch">Boosting e-commerce search by profit and popularity with the function score query in Elasticsearch</a>, and (2) <a href="https://www.elastic.co/search-labs/blog/ecommerce-search-relevance-cohort-aware-ranking-elasticsearch">How to improve e-commerce search relevance with personalized cohort-aware ranking</a>. Both approaches rely on the idea that business signals should guide BM25, not override it. Multiplicative boosting through <code>function_score</code> provides a practical, transparent, and scalable method for achieving that balance in real-world e-commerce search.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/bm25-ranking-multiplicative-boosting-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/bm25-ranking-multiplicative-boosting-elasticsearch</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Alexander Marquardt]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt72914b2406032448/6a170e2c60084b4e913c45f9/6150bb846170d9be926a19260846a161ed377a5f-1098x542.png" length="0" type="image/png"/>
    <pubDate>Mon, 22 Dec 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch Serverless pricing demystified: VCUs and ECUs explained]]></title>
    <description><![CDATA[Learn how Elasticsearch Serverless pricing works for Elastic’s fully-managed deployment offering. We explain VCUs (Search, Ingest, ML) and ECUs, detailing how consumption is based on actual allocated resources, workload complexity, and Search Power.]]></description>
    <content:encoded><![CDATA[<p><em>Navigating Elasticsearch Serverless pricing is simple... you pay for the resources you use. Getting a handle on VCUs, ECUs, and the factors that drive your consumption is key to making informed decisions about your usage. In this blog, we'll break down exactly how Elasticsearch Serverless pricing works so you can plan, monitor, and optimize your spend.</em></p><p>When we built Elasticsearch Serverless, we had to decide how to bill our users. While a charge per query may have been easier to reason about from a consumption perspective, it would be a lot harder to reason about from a resource perspective. Instead, we implemented a simple pricing scheme comprising three dimensions for compute: search, ingest, and machine learning VCUs. This means we charge users for the actual resources we allocate to fulfill your requested workloads.</p><h2>VCU, ECU, and other terms</h2><p>Let's start by defining a few terms that will keep coming back throughout this post.</p><h3>VCU</h3><p>A VCU is a <a href="https://www.elastic.co/docs/deploy-manage/cloud-organization/billing/elasticsearch-billing-dimensions#elasticsearch-billing-information-about-the-vcu-types-search-ingest-and-ml">Virtual Compute Unit</a>, representing a fraction of RAM, CPU, and local disk for caching. We separate compute by the workloads they support, so we have three flavors of VCU:</p><ol><li><p>Search VCU</p></li><li><p>Ingest VCU</p></li><li><p>Machine Learning (ML) VCU</p></li></ol><p>VCU’s are charged by the hour.</p><h3>Regional pricing</h3><p>We have different prices for different regions and different cloud providers. You can find a full list of prices <a href="https://cloud.elastic.co/cloud-pricing-table?productType=serverless">on this page</a>.</p><h3>ECU</h3><p>An ECU is an <a href="https://www.elastic.co/docs/deploy-manage/cloud-organization/billing/ecu">Elastic Consumption Unit</a>, which is the unit we bill you in. The nominal value of an ECU is $1.00 USD. All of the different components of consumption are charged at a specific rate of ECUs per time unit. For example, one Gigabyte of storage might cost 0.047 ECU per month, so 100 GB of storage will cost you 4.7 ECU = $4.70 for one month. Similarly, if your search workload consumed 10 VCUs in a day and the Search VCU rate in your region is 0.09 ECU, your cost for that day would be $0.90.</p><h3>Interactive Dataset Size</h3><p>The amount of data in your project has a direct influence on your costs. We make the distinction of “interactive dataset” primarily for time-series data, as this relates to the amount of data in the Boost Window. For non-time-series data, this is simply the amount of data in the project.</p><p></p><h2>Project settings</h2><p>We have three <a href="https://www.elastic.co/docs/deploy-manage/deploy/elastic-cloud/project-settings">project settings</a> that allow you to control your project's usage.</p><h3>Search power</h3><p>Search Power controls the speed of searches against your data. With Search Power, you can improve search performance by adding more resources for querying, or you can reduce provisioned resources to cut costs. Choose from three Search Power settings:</p><p><strong>On-demand</strong>: Autoscales based on data and search load, with a lower minimum baseline for resource use. This flexibility results in more variable query latency and reduced maximum throughput.</p><p><strong>Performant</strong>: Delivers consistently low latency and autoscales to accommodate moderately high query throughput.</p><p><strong>High-availability</strong>: Optimized for high-throughput scenarios, autoscaling to maintain query latency even at very high query volumes.</p><h3>Boost window</h3><p>For time series use cases, the boost window is the number of days of data that constitutes your interactive dataset size. The interactive dataset is the portion of your data that we keep cached, and that we use to determine how to scale the Search tier for your project. By default, the boost window is seven days.</p><h3>Data retention</h3><p>You can set the number of days of data that are retained in your project, which will affect the amount of storage we need. You can do this on a per-data stream basis in your project.</p><h2>Price components</h2><p>Serverless Elasticsearch contains a few different pricing components. For most use cases, the components you will care most about are Search, Ingest, and ML VCUs, as well as the Elastic Inference Service's token consumption.</p><h3>Search VCUs</h3><p>Search VCU consumption is the most complex part of pricing. We make this simple for you by automatically determining the right amount of VCUs that are needed to fulfill your workloads. For more details on how our autoscaling logic works, see <a href="https://www.elastic.co/search-labs/blog/elasticsearch-serverless-tier-autoscaling">our earlier blog on the topic</a>.</p><h4>Search VCU inputs</h4><p>Search VCUs are allocated based on a few factors, but mainly, we can boil it down to three inputs: the interactive dataset size, the search load on the system, and Search Power.</p><p>For traditional search use cases, the interactive dataset size will generally be your entire dataset. For time series use cases, it will be the portion of your dataset that fits inside the Boost Window.</p><p>Search load measures the amount of load being placed on the system by currently active searches. The main contributing factors are the number of searches per second, the complexity of the searches (the more that needs to be computed, the higher the load), and the size of the dataset that needs to be searched to fulfill the result. If we can get you the right number of results by scanning 10% of the dataset, then the load will be much lower than if we need to scan the full dataset.</p><p>Finally, <a href="https://www.elastic.co/docs/deploy-manage/deploy/elastic-cloud/project-settings">Search Power</a> influences the number of VCUs we allocate. Each Search Power setting defines the baseline capacity of the search tier.</p><p>In short: the larger the dataset size and the higher the search load, the more VCUs we need to fulfill your search requests. Search Power allows you to tune to what extent we will scale up and down.</p><h4>Minimum VCUs</h4><p>Elasticsearch Serverless is designed to align infrastructure costs directly with your application's demand. </p><p>For smaller workloads, the search infrastructure can scale down to zero VCUs during periods of inactivity. If the system detects fifteen minutes of total inactivity, the associated hardware resources are deprovisioned. This makes the platform highly cost-effective for development environments, bursty workloads, or applications with intermittent usage. Note that inactivity means actual inactivity: no user-initiated searches whatsoever. As soon as we need to serve a search of any kind, we need to allocate hardware resources to execute that search.</p><p>As your interactive dataset grows, the system eventually reaches a storage threshold where a baseline level of resources is required to maintain data availability and indexing readiness. A minimum VCU allocation is maintained to ensure your data remains "warm" and queryable, even if no active searches are occurring.</p><h4>VCU consumption is not linear</h4><p>Because our hardware is allocated in steps, consumption of VCUs does not necessarily scale linearly with workload size. Each scaling step can contain a wide range of workloads, and if your workload is at the bottom of that range, it may have a lot of room to grow before we need to jump to the next scaling step.</p><p>This can make estimating based on a non-representative workload hard. For example, you may be consuming 2 VCUs per hour on a small workload. It's entirely possible that you could increase your workload size by a factor of 100 and still fit in that 2 VCU per hour load before we need to start increasing the amount of VCUs we allocate to serve your workload.</p><p>We know this makes estimating your cost a little harder, and we are working on ways to make that easier for you. If you need more help estimating your likely price, you can always talk to our customer team and get more personalized assistance.</p><h2>Ingest VCUs</h2><p>Ingest VCUs are much simpler than Search VCUs.</p><h4>Ingest VCU Inputs</h4><p>Ingest VCUs have essentially three inputs: the number of indices, the ingest rate, and the ingest complexity. We need to allocate a little bit of memory for every index in your system, which is why the number of indices matters. Read indices in data streams do not count for this calculation.</p><p>The faster you ingest, the more CPU we will need to process that ingestion. And the more complex your ingest requests, the more CPU we will need. Some factors that make ingest requests more expensive to execute are complicated field mappings or a lot of post-processing.</p><h4>Minimum Ingest VCUs</h4><p>We do not have a minimum number of VCUs we allocate to your ingest. If you do not ingest data, we do not need to allocate any VCUs to processing ingestion. There is an exception for a large number of indices (think: thousands of indices), where we do need to keep some resources allocated to be responsive when indexing requests come in.</p><h4>VCU consumption is not linear</h4><p>As with Search VCUs, we allocate Ingest VCUs based on step functions. Each step can contain a wide range of workloads: it's entirely possible that if you have a minimal amount of ingest, you could increase your ingest rate by a factor of 100 and still fit in the same step, thus not actually increasing your cost.</p><h2>AI workloads</h2><p>When running machine learning tasks in Serverless, we give you three options:</p><ol><li><p>You use our Elastic Inference Service (EIS) to run your inference and completion workloads. We take care of everything, and you are charged per token.</p></li><li><p>You use traditional Elasticsearch Machine Learning capabilities to run your workloads. These use our Trained Models capabilities. We will scale up and down based on your machine learning workload requirements.</p></li><li><p>You do it yourself, outside of our systems, and just bring your vectors or other inference results to store and search in Elasticsearch.</p></li></ol><h4>EIS</h4><p>The pricing for EIS is <a href="https://cloud.elastic.co/cloud-pricing-table?productType=serverless">quite straightforward</a>: you get charged a rate per one million consumed tokens. Token consumption is generally easy to predict for inference workloads. For LLM-based tasks, particularly agentic ones, this can be more complex, and some experimentation and trial runs may be useful to determine how many tokens your workloads typically consume.</p><h4>ML VCUs</h4><p>Machine Learning VCUs work on one simple input: machine learning workloads. The more inference you require, the more VCUs we will consume. Once you stop performing inference, we will scale down. We will keep a trained model in memory for about 24 hours after you last used it so that we can be responsive, which means that the minimal amount of VCU required to keep that model available will remain up for 24 hours before scaling down entirely.</p><p>We generally recommend our customers use EIS instead of our Machine Learning nodes for inference, particularly if your usage is periodic. By switching to EIS, you will not have to wait for machine learning nodes to spin up, and we won't charge you for unused ML node time before scaling down. EIS charges on a per token basis.</p><h2>Storage</h2><p>We charge storage per gigabyte per month. Storage does serve as an input into other parts of our system, particularly Search VCUs (see Search VCU above), but the pricing for storage itself is <a href="https://cloud.elastic.co/cloud-pricing-table?productType=serverless">quite straightforward</a>.</p><h2>Data Out (egress)</h2><p>We charge you for the data you take out of the system.</p><p>To minimize your egress costs, we recommend a few optimizations on your queries:</p><ol><li><p>Do not return vectors in your query responses. We <a href="https://www.elastic.co/search-labs/blog/elasticsearch-exclude-vectors-from-source">do this by default</a> for indices created after October 2025. You can always return vectors in your responses explicitly if necessary.</p></li><li><p>Return only the fields needed for your application. You can <a href="https://www.elastic.co/search-labs/blog/displaying-fields-in-an-elasticsearch-index">do this</a> by using the <code>fields</code> and <code>_source</code> parameters.</p></li></ol><h2>Support</h2><p>We charge <a href="https://www.elastic.co/pricing/serverless-search">support</a> as a percentage of your total ECU usage. We currently have four levels of support:</p><ol><li><p>Limited support</p></li><li><p>Base support</p></li><li><p>Enhanced support</p></li><li><p>Premium support</p></li></ol><h2>Project subtype profiles</h2><p>We currently offer two project subtypes for Serverless Elasticsearch, referred to as “General Purpose” and “Vector Optimized”. All Serverless Elasticsearch projects created through the cloud console UI will be created using the “General Purpose” option. You may create a “Vector Optimized” by calling the API directly with the <code>optimized_for</code> parameter (see <a href="https://www.elastic.co/docs/api/doc/elastic-cloud-serverless/operation/operation-createelasticsearchproject">documentation</a> for all options).</p><p>The difference between the two options is the allocation of resources. We allocate approximately four times more resources (aka VCUs) to the “Vector Optimized” profile, which will result in your costs being up to four times higher. This is why we recommend starting on the “General Purpose” profile and only using the “Vector Optimized” profile when your use case demands the use of uncompressed dense vectors with high dimensionality, and quantization and <a href="https://www.elastic.co/search-labs/blog/diskbbq-elasticsearch-introduction">DiskBBQ</a> will not serve your needs.</p><p>When Serverless Elasticsearch was envisioned years ago, we thought that vector workloads would require much more resources to remain performant. However, with innovations like <code>semantic_text</code>, <code>sparse_vector</code> models, and <a href="https://www.elastic.co/search-labs/blog/elasticsearch-9-1-bbq-acorn-vector-search">Better Binary Quantization</a> (BBQ), we’ve found that many vector workloads perform well on the “General Purpose” profile at a fraction of the cost. Therefore, don’t let the “Vector Optimized” label fool you…you can get excellent price <em>and</em> performance for vector workloads on the “General Purpose” profile.</p><h2>Monitoring costs</h2><p>We recognize that keeping track of your costs, especially when you are new to Elasticsearch Serverless, is important to you. We built a few tools just for this purpose, and continue to improve them for even greater visibility.</p><h2>Cloud console billing usage</h2><p>The <a href="https://www.elastic.co/docs/deploy-manage/cloud-organization/billing/view-billing-history">Elastic Cloud Console</a> provides billing details for your cloud account, across all cloud-based resources, including Elasticsearch Serverless. There, you can find a breakdown of all the price components described in this article. Filters allow you to zoom in on specific time periods and resources.</p><p>To further monitor your costs, you can also configure custom <a href="https://www.elastic.co/docs/deploy-manage/cloud-organization/billing/manage-billing-notifications">budget alerts </a>from the Budgets and notifications tab under the Billing and subscriptions page.</p><h2>AutoOps monitoring</h2><p>We’re bringing <a href="https://www.elastic.co/docs/deploy-manage/monitor/autoops/autoops-for-serverless">AutoOps to Serverless</a>! One of the key value propositions of Elasticsearch Serverless is that we ensure everything runs smoothly, but that also means you have limited observability into the infrastructure. AutoOps for Serverless gives users visibility into what is driving usage, and, therefore, costs.</p><p>AutoOps is rolled out in new Serverless regions regularly, and we're always working to add new monitoring tools. Make sure to check out the <a href="https://www.elastic.co/docs/deploy-manage/monitor/autoops/ec-autoops-regions#autoops-for-serverless-full-regions">region coverage</a> and future planned monitoring tools.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-serverless-pricing-vcus-ecus</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-serverless-pricing-vcus-ecus</guid>
    <category><![CDATA[Basics]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Sander Philipse,Pete Galeotti]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3b8542204a8988dc/6a170bed0e2e49cd2641a12e/46f1e3c09e17cb8aa2a1cca64624bf533e55fe1d-1746x1096.png" length="0" type="image/png"/>
    <pubDate>Fri, 19 Dec 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Evaluating search query relevance with judgment lists]]></title>
    <description><![CDATA[Explore how to build judgment lists to objectively evaluate search query relevance and improve performance metrics such as recall, for scalable search testing in Elasticsearch.]]></description>
    <content:encoded><![CDATA[<p>Developers working on search engines often encounter the same issue: the business team is not satisfied with one particular search because the documents they expect to be at the top of the search results appear third or fourth on the list of results.</p><p>However, when you fix this one issue, you accidentally break other queries since you couldn’t test all cases manually. But how can you or your QA team test if a change in one query has a ripple effect in other queries? Or even more importantly, how can you be sure that your changes actually improved a query?</p><h2>Towards a systematic evaluation</h2><p>Here is where judgment lists come in useful. Instead of depending on manual and subjective testing any time you make a change, you can define a fixed set of queries that are relevant for your business case, together with their relevant results.</p><p>This set becomes your baseline. Every time you implement a change, you use it to evaluate if your search actually improved or not.</p><p>The value of this approach is that it:</p><ul><li><p><strong>Removes uncertainty</strong>: you no longer need to wonder if your changes impact other queries; the data will tell you.</p></li><li><p><strong>Stops manual testing</strong>: once the judgment sets are recorded, the test is automatic.</p></li><li><p><strong>Supports changes</strong>: You can show clear metrics that support the benefits of a change.</p></li></ul><h2>How to start building your judgment list</h2><p>One of the easiest ways to start is to take a representative query and manually select the relevant documents. There are two ways to do this list:</p><ul><li><p><strong>Binary Judgments:</strong> Each document associated with a query gets a <strong>simple tag</strong>: <em>relevant</em> (usually with a score of “1”) and not-relevant (“0”).</p></li><li><p><strong>Graded Judgments:</strong> Here, each document gets a score with different levels. For example: setting a 0 to 4 scale, similar to a <a href="https://en.wikipedia.org/wiki/Likert_scale">Likert scale</a>, where 0 = “not at all relevant” and 4 = “totally relevant,” with variations like “relevant,” “somewhat relevant,” etc.</p></li></ul><p>Binary judgments work well when the search intent has clear limits: Should this document be in the results or not?</p><p>Graded judgements are more useful when there are grey areas: some results are better than others, so you can get “very good,” “good,” and “useless” results and use metrics that value the order of the results and the user’s feedback. However, graded scales also introduce drawbacks: different reviewers may use the scoring levels differently, which makes the judgments less consistent. And because graded metrics give more weight to higher scores, even a small change (like rating something a 3 instead of a 4) can create a much bigger shift in the metric than the reviewer intended. This added subjectivity makes graded judgments noisier and harder to manage over time.</p><h2>Do I need to classify the documents myself?</h2><p>Not necessarily, since there are different ways to create your judgment list, each with its own advantages and disadvantages:</p><ul><li><p><strong>Explicit Judgments:</strong> Here, SMEs go over each query/document and manually decide if (or how) relevant it is. Though this provides quality and control, it is less scalable.</p></li><li><p><strong>Implicit Judgments:</strong> With this method, you infer the relevant documents based on real-user behavior like clicks, bounce rate, and purchases, among others. This approach allows you to gather data automatically, but it might be biased. For example, users tend to click top results more often, even if they are not relevant.</p></li><li><p><strong>AI-Generated Judgments:</strong> This last option uses models (like LLMs) to automatically evaluate queries and documents, often referred to as <a href="https://en.wikipedia.org/wiki/LLM-as-a-Judge">LLM juries</a>. It’s fast and easy to scale, but the quality of the data depends on the quality of the model you’re using and how well LLM training data aligns with your business <a href="http://interests.as/">interests</a>. As with human grades, LLM juries can introduce their own biases or inconsistencies, so it’s important to validate their output against a smaller set of trusted judgments. LLM models are probabilistic by nature, so it is not uncommon to see an LLM model giving different grades to the same result regardless of setting <a href="https://www.ibm.com/think/topics/llm-temperature">temperature</a> parameter as 0.</p></li></ul><p>Below are some recommendations to choose the best method for creating your judgment set:</p><ul><li><p>Decide how critical some features are for you that only users can properly judge (like price, brand, language, style, and product details). If those are critical, you need <strong>explicit judgments</strong> for at least some part of your <em>judgment list</em>.</p></li><li><p>Use <strong>implicit judgements</strong> when your search engine already has enough traffic so you can use clicks, conversions, and lingering time metrics to detect usage trends. You should still interpret these carefully, contrasting them with your explicit judgement sets to prevent any bias (e.g: users tend to click top-ranked results more often, even if lower-ranked results are more relevant)</p></li></ul><p>To address this, position debiasing techniques adjust or reweight click data to better reflect true user interest. Some approaches include:</p><ul><li><p><strong>Results shuffling</strong>: Change the order of search results for a subset of users to estimate how position affects clicks.</p></li><li><p><strong>Click models </strong>include<a href="https://wiki.math.uwaterloo.ca/statwiki/index.php?title=a_Dynamic_Bayesian_Network_Click_Model_for_web_search_ranking">Dynamic Bayesian Network </a><a href="https://wiki.math.uwaterloo.ca/statwiki/index.php?title=a_Dynamic_Bayesian_Network_Click_Model_for_web_search_ranking"><strong>DBN</strong></a>, <a href="https://rsrikant.com/papers/kdd10.pdf">User Browsing Model </a><a href="https://rsrikant.com/papers/kdd10.pdf"><strong>UBM</strong></a>. These Statistical models estimate the probability of a click reflects real interest rather than just position, using patterns like scrolling, dwell time, click sequence, and returning to the results page.</p></li></ul><h2>Example: Movie rating app</h2><h3>Prerequisites</h3><p>To run this example, you need a running Elasticsearch 8.x cluster, <a href="https://www.elastic.co/downloads/elasticsearch">locally</a> or <a href="https://www.elastic.co/cloud/cloud-trial-overview">Elastic Cloud</a> (Hosted or Serverless), and access to the <a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis">REST API</a> or Kibana.</p><p>Think about an app in which users can upload their opinions about movies and also search for movies to watch. As the texts are written by users themselves, they can have typos and many variations in terms of expression. So it’s essential that the search engine is able to interpret that diversity and provide helpful results for the users.</p><p>To be able to iterate queries without impacting the overall search behavior, the business team in your company created the following binary judgment set, based on the most frequent searches:</p><p>Query</p><p>DocID</p><p>Text</p><p>DiCaprio performance</p><p>doc1</p><p>DiCaprio's performance in The Revenant was breathtaking.</p><p>DiCaprio performance</p><p>doc2</p><p>Inception shows Leonardo DiCaprio in one of his most iconic roles.</p><p>DiCaprio performance</p><p>doc3</p><p>Brad Pitt delivers a solid performance in this crime thriller.</p><p>DiCaprio performance</p><p>doc4</p><p>An action-packed adventure with stunning visual effects.</p><p>sad movies that make you cry</p><p>doc5</p><p>A heartbreaking story of love and loss that made me cry for hours.</p><p>sad movies that make you cry</p><p>doc6</p><p>One of the saddest movies ever made — bring tissues!</p><p>sad movies that make you cry</p><p>doc7</p><p>A lighthearted comedy that will make you laugh</p><p>sad movies that make you cry</p><p>doc8</p><p>A science-fiction epic full of action and excitement.</p><p>Creating the index:</p>PUT movies
{
  "mappings": {
    "properties": {
      "text": {
        "type": "text"
      }
    }
  }
}<p>BULK request:</p>POST /movies/_bulk
{ "index": { "_id": "doc1" } }
{ "text": "DiCaprio performance in The Revenant was breathtaking." }
{ "index": { "_id": "doc2" } }
{ "text": "Inception shows Leonardo DiCaprio in one of his most iconic roles." }
{ "index": { "_id": "doc3" } }
{ "text": "Brad Pitt delivers a solid performance in this crime thriller." }
{ "index": { "_id": "doc4" } }
{ "text": "An action-packed adventure with stunning visual effects." }
{ "index": { "_id": "doc5" } }
{ "text": "A heartbreaking story of love and loss that made me cry for hours." }
{ "index": { "_id": "doc6" } }
{ "text": "One of the saddest movies ever made -- bring tissues!" }
{ "index": { "_id": "doc7" } }
{ "text": "A lighthearted comedy that will make you laugh." }
{ "index": { "_id": "doc8" } }
{ "text": "A science-fiction epic full of action and excitement." }<p>Below is the Elasticsearch query the app is using:</p>GET movies/_search
{
 "query": {
   "match": {
     "text": {
       "query": "DiCaprio performance",
       "minimum_should_match": "100%"
     }
   }
 }
}<h3>From judgment to metrics</h3><p>By themselves, judgment lists do not provide much information; they are only an expectation of the results from our queries. Where they really shine is when we use them to calculate objective metrics to measure our search performance.</p><p>Nowadays, most of the popular metrics include</p><ul><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-rank-eval#k-precision"><strong>Precision</strong></a><strong>: </strong>Measures the proportion of results that are truly relevant within all search results.</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-rank-eval#k-recall"><strong>Recall</strong></a><strong>: </strong>Measures the proportion of relevant results the search engine found among x results.</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-rank-eval#_discounted_cumulative_gain_dcg"><strong>Discounted Cumulative Gain (DCG)</strong></a><strong>: </strong>Measures the quality of the result’s ranking, considering the most relevant results should be at the top.</p></li><li><p><a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-rank-eval#_mean_reciprocal_rank"><strong>Mean Reciprocal Rank (MRR):</strong></a> Measures the position of the first relevant result. The higher it is in the list, the higher its score.</p></li></ul><p>Using the same movie rating app as an example, we’ll calculate the recall metric to see if there’s any information that is being left out of our queries.</p><p>In Elasticsearch, we can use the <em>judgment lists</em> to calculate metrics via the <a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-rank-eval">Ranking Evaluation API</a>. This API receives as input the judgment list, the query, and the metric you want to evaluate, and returns a value, which is a comparison of the query result with the judgment list.</p><p>Let’s run the judgment list for the two queries that we have:</p>POST /movies/_rank_eval
{
 "requests": [
   {
     "id": "dicaprio-performance",
     "request": {
       "query": {
         "match": {
           "text": {
             "query": "DiCaprio performance",
             "minimum_should_match": "100%"
           }
         }
       }
     },
     "ratings": [
       {
         "_index": "movies",
         "_id": "doc1",
         "rating": 1
       },
       {
         "_index": "movies",
         "_id": "doc2",
         "rating": 1
       },
       {
         "_index": "movies",
         "_id": "doc3",
         "rating": 0
       },
       {
         "_index": "movies",
         "_id": "doc4",
         "rating": 0
       }
     ]
   },
   {
     "id": "sad-movies",
     "request": {
       "query": {
         "match": {
           "text": {
             "query": "sad movies that make you cry",
             "minimum_should_match": "100%"
           }
         }
       }
     },
     "ratings": [
       {
         "_index": "movies",
         "_id": "doc5",
         "rating": 1
       },
       {
         "_index": "movies",
         "_id": "doc6",
         "rating": 1
       },
       {
         "_index": "movies",
         "_id": "doc7",
         "rating": 0
       },
       {
         "_index": "movies",
         "_id": "doc8",
         "rating": 0
       }
     ]
   }
 ],
 "metric": {
   "recall": {
     "k": 10,
     "relevant_rating_threshold": 1
     }
 }
}<p>We’ll use two requests to _rank_eval: one for the DiCaprio query and another for sad movies. Each request includes a query and its judgment list (ratings). We don’t need to grade all documents since the ones not included in the ratings are considered as with no judgment. To do the calculations, recall only considers the “relevant set,” the documents that are considered relevant in the rating.</p><p>In this case, the DiCaprio query has a recall of 1, while the sad movies got 0. This means that for the first query, we were able to get all relevant results, while in the second query, we did not get any. The average recall is therefore 0.5.</p>{
 "metric_score": 0.5,
 "details": {
   "dicaprio-performance": {
     "metric_score": 1,
     "unrated_docs": [],
     "hits": [
       {
         "hit": {
           "_index": "movies",
           "_id": "doc1",
           "_score": 2.4826927
         },
         "rating": 1
       },
       {
         "hit": {
           "_index": "movies",
           "_id": "doc2",
           "_score": 2.0780432
         },
         "rating": 1
       }
     ],
     "metric_details": {
       "recall": {
         "relevant_docs_retrieved": 2,
         "relevant_docs": 2
       }
     }
   },
   "sad-movies": {
     "metric_score": 0,
     "unrated_docs": [],
     "hits": [],
     "metric_details": {
       "recall": {
         "relevant_docs_retrieved": 0,
         "relevant_docs": 2
       }
     }
   }
 },
 "failures": {}
}<p>Maybe we’re being too strict with the <strong>minimum_should_match </strong>parameter since by demanding that 100% of the words in the query are found in the documents, we’re probably leaving relevant results out. Let’s remove the <strong>minimum_should_match</strong> parameter so that a document is considered relevant if only one word in the query is found in it.</p>POST /movies/_rank_eval
{
 "requests": [
   {
     "id": "dicaprio-performance",
     "request": {
       "query": {
         "match": {
           "text": {
             "query": "DiCaprio performance"
           }
         }
       }
     },
     "ratings": [
       {
         "_index": "movies",
         "_id": "doc1",
         "rating": 1
       },
       {
         "_index": "movies",
         "_id": "doc2",
         "rating": 1
       },
       {
         "_index": "movies",
         "_id": "doc3",
         "rating": 0
       },
       {
         "_index": "movies",
         "_id": "doc4",
         "rating": 0
       }
     ]
   },
   {
     "id": "sad-movies",
     "request": {
       "query": {
         "match": {
           "text": {
             "query": "sad movies that make you cry"
           }
         }
       }
     },
     "ratings": [
       {
         "_index": "movies",
         "_id": "doc5",
         "rating": 1
       },
       {
         "_index": "movies",
         "_id": "doc6",
         "rating": 1
       },
       {
         "_index": "movies",
         "_id": "doc7",
         "rating": 0
       },
       {
         "_index": "movies",
         "_id": "doc8",
         "rating": 0
       }
     ]
   }
 ],
 "metric": {
   "recall": {
     "k": 10,
     "relevant_rating_threshold": 1
     }
 }
}<p>As you can see, by removing the <strong>minimum_should_match</strong> parameter in one of the two queries, we now get an average recall of 1 in both.</p>{
  "metric_score": 1,
  "details": {
    "dicaprio-performance": {
      "metric_score": 1,
      "unrated_docs": [],
      "hits": [
        {
          "hit": {
            "_index": "movies",
            "_id": "doc1",
            "_score": 2.0661702
          },
          "rating": 1
        },
        {
          "hit": {
            "_index": "movies",
            "_id": "doc3",
            "_score": 0.732218
          },
          "rating": 0
        },
        {
          "hit": {
            "_index": "movies",
            "_id": "doc2",
            "_score": 0.6271719
          },
          "rating": 1
        }
      ],
      "metric_details": {
        "recall": {
          "relevant_docs_retrieved": 2,
          "relevant_docs": 2
        }
      }
    },
    "sad-movies": {
      "metric_score": 1,
      "unrated_docs": [],
      "hits": [
        {
          "hit": {
            "_index": "movies",
            "_id": "doc7",
            "_score": 2.1307156
          },
          "rating": 0
        },
        {
          "hit": {
            "_index": "movies",
            "_id": "doc5",
            "_score": 1.3160692
          },
          "rating": 1
        },
        {
          "hit": {
            "_index": "movies",
            "_id": "doc6",
            "_score": 1.190063
          },
          "rating": 1
        }
      ],
      "metric_details": {
        "recall": {
          "relevant_docs_retrieved": 2,
          "relevant_docs": 2
        }
      }
    }
  },
  "failures": {}
}<p>In summary, removing the minimum_should_match: 100% clause, allows us to got a perfect recall for both queries.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltaf4f08a8a2915180/6a170df61949f76cbfe7aaba/24d055da4348c63827ba7046fe8cafb6f47cadd8-546x628.png" alt="" /><p>We did it! Right?</p><p>Not so fast!</p><p>By improving recall, we open the door to a wider range of results. However, each adjustment implies a trade-off. This is why defining complete test cases, using different metrics to evaluate changes.</p><p>Using judgment lists and metrics prevents you from going in blind when making changes since you now have data to back them up. Validation is no longer manual and repetitive, and you can test your changes in more than just one use case. Additionally, A/B testing allows you to test live which configuration works best for your users and business case, thus coming full circle from technical metrics and real-world metrics.</p><h2>Final recommendations for using judgment lists</h2><p>Working with judgment lists is not only about measuring but also about creating a framework that allows you to iterate with confidence. To achieve this, you can follow these recommendations:</p><ol><li><p><strong>Start small, but start</strong>. You don’t need to have 10,000 queries with 50 judgment lists each. You only need to identify the 5–10 most critical queries for your business case and define which documents you expect to see at the top of the results. This already gives you a base. You typically want to start with the top queries plus the queries with no results. You can also start testing with an easy-to-configure metric like Precision and then work your way up in complexity.</p></li><li><p><strong>Validate with users.</strong> Complement the numbers with A/B testing in production. This way, you’ll know if changes that look good in the metrics are also generating a real impact.</p></li><li><p><strong>Keep the list alive.</strong> Your business case will evolve, and so will your critical queries. Update your judgment periodically to reflect new needs.</p></li><li><p><strong>Make it part of the flow.</strong> Integrate judgment lists into your development pipelines. Make sure each configuration change, synonym, or text analysis is automatically validated against your base list.</p></li><li><p><strong>Connect technical knowledge with strategy.</strong> Don’t stop at measuring technical metrics like precision or recall. Use your evaluation results to inform business outcomes.</p></li></ol>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/judgment-lists-search-query-relevance-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/judgment-lists-search-query-relevance-elasticsearch</guid>
    <category><![CDATA[Relevance]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Jhon Guzmán]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcadfd2fb1cc95b4c/6a170df7acf0887798be9bd0/25478d0ffb228afd5d65d82312998ec1c299c565-700x490.png" length="0" type="image/png"/>
    <pubDate>Thu, 11 Dec 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[GenAI for customer support — Part 4: Tuning RAG search for relevance]]></title>
    <description><![CDATA[This series gives you an inside look at how we're using generative AI in customer support. Join us as we share our journey in real-time, focusing in this section on tuning RAG search for relevance.]]></description>
    <content:encoded><![CDATA[<p></p>This blog series reveals how our Field Engineering team used the Elastic stack with generative AI to develop a lovable and effective customer support chatbot. If you missed other installments in the series, be sure to check out <a href="https://www.elastic.co/blog/genai-customer-support-building-proof-of-concept">part one</a>, <a href="https://search-labs.elastic.co/search-labs/blog/genai-customer-support-building-a-knowledge-library">part two</a>, <a href="https://search-labs.elastic.co/search-labs/blog/genai-elastic-elser-chat-interface">part three</a>, the <a href="https://www.elastic.co/blog/generative-ai-customer-support-elastic-support-assistant">launch blog</a>, and <a href="https://www.elastic.co/search-labs/blog/genai-customer-support-observability">part five</a>.<p>
Welcome to part 4 of our blog series on integrating generative AI in Elastic's customer support. This installment dives deep into the role of Retrieval-Augmented Generation (RAG) in enhancing our AI-driven Technical Support Assistant. Here, we address the challenges, solutions, and outcomes of refining search effectiveness, providing action items to further improve its capabilities using the toolset provided in the Elastic Stack version <em>8.11</em>.</p><p>Implied by those actions, we have achieved a <strong>~75% increase in top-3 results</strong> relevance and gained over <strong>300,000 AI-generated summaries that we can leverage for all kinds of future applications</strong>. If you're new to this series, be sure to review the earlier posts that introduce the core technology and architectural setup. If you missed the last blog of the series, you can find it <a href="https://www.elastic.co/search-labs/blog/genai-elastic-elser-chat-interface">here</a>.</p><h2>RAG tuning: A search problem</h2><p>Perfecting RAG (Retrieval-Augmented Generation) is fundamentally about hitting the bullseye in search accuracy 🎯:</p><ul><li><p>Like an archer carefully aiming to hit the center of the target, we want to focus on <strong>accuracy</strong> for each hit.</p></li><li><p>Not only that, we also want to ensure that we have the best targets to hit – or <strong>high-quality data</strong>.</p></li></ul><p>Without <strong>both together</strong>, there's the potential risk that large language models (LLMs) might hallucinate and generate misleading responses. Such mistakes can definitely shake users' trust in our system, leading to a deflecting usage and poor return on investment.</p><p>To avoid those negative implications, we've encountered several challenges that have helped us refine our search accuracy and data quality over the course of our journey. These challenges have been instrumental in shaping our approach to tuning RAG for relevance, and we're excited to share our insights with you.</p><p>That said: <strong>let's dive into the details!</strong></p><h2>Our first approach</h2><p>We started with a lean, effective solution that could quickly get us a valuable RAG-powered chatbot in production. This meant focusing on key functional aspects that would bring it to operational readiness with optimal search capabilities. To get us into context, we'll make a quick walkthrough around four key vital components of the Support AI Assistant: <a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#data"><em>data</em></a>, <a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#query"><em>querying</em></a>, <a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#generation"><em>generation</em></a>, and <a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#feedback"><em>feedback</em></a>.</p><h3>Data</h3><p>As showcased in the <a href="https://search-labs.elastic.co/search-labs/blog/genai-customer-support-building-a-knowledge-library#elastic-supports-knowledge-library">2nd blog article of this series</a>, our journey began with an extensive database that included over 300,000 documents consisting of <em>Technical Support Knowledge Articles</em> and various pages crawled from our website, such as Elastic's <em>Product Documentation</em> and <em>Blogs</em>. This rich dataset served as the foundation for our search queries, ensuring a broad spectrum of information about Elastic products was available for precise retrieval. To this end, we leveraged Elasticsearch to store and search our data.</p><h3>Query</h3><p>Having great <a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#data">data</a> to search by, it's time to talk about our querying component. We adopted a standard Hybrid-Search strategy, which combines the traditional strengths of <strong>BM25</strong>, <em>Keyword-based Search</em>, with the capabilities of <em>Semantic Search</em>, powered by <strong>ELSER</strong>.</p><p>For the semantic search component, we used <code>text_expansion</code> queries against both <code>title</code> and <code>summary</code> embeddings. On the other hand, for broad keyword relevance we search multiple fields using <code>cross_fields</code>, with a <code>minimum_should_match</code> parameter tuned to better perform with longer queries. Phrase matches, which often signal greater relevance, receive a higher boost. Here’s our initial setup:</p>const searchResults = await client.elasticsearchClient({
  // Alias pointing to the knowledge base indices.
  index: "knowledge-search", 
  body: {
    size: 3,
    query: {
      bool: {
        should: [
          // Keyword-search Component. 
          {
            multi_match: {
              query,
              // For queries with 3+ words, at least 49% must match.
              minimum_should_match: "1&lt;-1 3&lt;49%", 
              type: "cross_fields",
              fields: [
                "title",
                "summary",
                "body",
                "id",
              ],
            },
          },
          {
            multi_match: {
              query,
              type: "phrase",
              boost: 9,
              fields: [
                // Stem-based versions of our fields. 
                "title.stem",
                "summary.stem",
                "body.stem",
              ],
            },
          },
          // Semantic Search Component.
          {
            text_expansion: {
              "ml.inference.title_expanded.predicted_value": {
                model_id: ".elser_model_2",
                model_text: query,
              },
            },
          },
          {
            text_expansion: {
              "ml.inference.summary_expanded.predicted_value": {
                model_id: ".elser_model_2",
                model_text: query,
              },
            },
          },
        ],
      },
    },
  },
});
<h3>Generation</h3><p>After search, we build up the system prompt with different sets of instructions, also contemplating the <em><strong>top 3</strong></em> search results as context to be used. Finally, we feed the conversation alongside the built context into the LLM, generating a response. Here's the pseudocode showing the described behavior:</p>// We then feed the context into the LLM, generating a response.
const { stopGeneration } = fetchChatCompletionAPI(
  {
    // The system prompt + vector search results.
    context: buildContext(searchResults), 
    // The entire conversation + the brand new user question.
    messages, 
    // Additional parameters.
    parameters: { model: LLM.GPT4 } 
  },
  {
    onGeneration: (event: StreamGenerationEvent) =&gt; {
      // Stream generation events back to the user interface here...
    }
  }
);
<p>The reason for not including more than 3 search results was the limited quantity of tokens available to work within our dedicated Azure OpenAI's GPT4 deployment (PTU), allied with a relatively large user base.</p><h3>Feedback</h3><p>We used a third-party tool to capture client-side events, connecting to <em>Big Query</em> for storage and making the JSON-encoded events accessible for comprehensive analysis by everyone on the team. Here's a glance into the Big Query syntax that builds up our feedback view. The</p><p><code>JSON_VALUE</code> function is a means to extract fields from the event payload:</p><p></p>  SELECT
    -- Extract relevant fields from event properties
    JSON_VALUE(event_properties, '$.chat_id') AS `Chat ID`,
    JSON_VALUE(event_properties, '$.input') AS `Input`,
    JSON_VALUE(event_properties, '$.output') AS `Output`,
    JSON_VALUE(event_properties, '$.context') AS `Context`,
    
    -- Determine the reaction (like or dislike) to the interaction
    CASE JSON_VALUE(event_properties, '$.reaction')
      WHEN 'disliked' THEN '👎'
      WHEN 'liked' THEN '👍'
    END AS `Reaction`,
    
    -- Extract feedback comment
    JSON_VALUE(event_properties, '$.comment') AS `Comment`,

    event_time AS `Time`
  FROM
    `frontend_events` -- Table containing event data
  WHERE
    event_type = "custom"
    AND JSON_VALUE(event_properties, '$.event_name') IN (
      'Chat Interaction', -- Input, output, context. 
      'Chat Feedback', -- Feedback comments.
      'Response Like/Dislike' -- Thumbs up/down.
    )
  ORDER BY `Chat ID` DESC, `Time` ASC; -- Order results by Chat ID and time
<p>We also took advantage of valuable direct feedback from internal users regarding the chatbot experience, enabling us to quickly identify areas where our search results did not match the user intent. Incorporating both would be instrumental in the discovery process that enabled us to refine our RAG implementation, as we're going to observe throughout the next section.</p><h2>Challenges</h2><p>With usage, interesting patterns started to emerge from feedback. Some user queries, like those involving specific <a href="https://en.wikipedia.org/wiki/Common_Vulnerabilities_and_Exposures"><em>CVEs</em></a> or <em>Product Versions</em> for instance, were yielding suboptimal results, indicating a disconnect between the user's intent and the <em>GenAI</em> responses. Let's take a closer look at the specific challenges identified, and how we solved them.</p><h3>#1: CVEs (Common Vulnerabilities and Exposures)</h3><p>Our customers frequently encounter alerts regarding lists of open CVEs that could impact their systems, often resulting in support cases. To address questions about those effectively, our dedicated internal teams meticulously maintain <em>CVE-type</em> Knowledge Articles. These articles provide standardized, official descriptions from Elastic, including detailed statements on the implications, and list the artifacts affected by each CVE.</p><p>Recognizing the potential of our chatbot to streamline access to this crucial information, our internal <em>InfoSec</em> and <em>Support Engineering</em> teams began exploring its capabilities with questions like this:</p>👨🏽 What are the implications of CVE's `2016-1837`, `2019-11756` and `2014-6439`?
<p>For such questions, one of the key advantages of using RAG – and also the main functional goal of adopting this design – is that we can pull up-to-date information, including it as context to the LLM and thus making it available instantly to produce awesome responses. That naturally will save us time and resources over fine-tuned LLM alternatives.</p><p>However, the produced responses wouldn't perform as expected. Essential to answer those questions, the search results often lacked relevance, a fact which we can confirm by looking closely at the search results for the example:</p>{
  ...
  "hits": [
    {
      "_index": "search-knowledge-articles",
      "_id": "...",
      "_score": 59.449028,
      "_source": {
        "id": "...",
        "title": "CVE-2019-11756", // Hit!
        "summary": "...",
        "body": "...",
        "category": "cve"
      }
    },
    {
      "_index": "search-knowledge-articles",
      "_id": "...",
      "_score": 42.15182,
      "_source": {
        "title": "CVE-2019-10172", // :(
        "summary": "...",
        "body": "...",
        "category": "cve"
      }
    },
    {
      "_index": "search-docs",
      "_id": "...",
      "_score": 38.413914,
      "_source": {
        "title": "Potential Sudo Privilege Escalation via CVE-2019-14287 | Elastic  Security Solution [8.11] | Elastic",  // :(
        "summary": "...",
        "body": "...",
        "category": "documentation"
      }
    }
  ]
}
<p>With just one relevant hit (<code>CVE-2019-10172</code>), we left the LLM without the necessary context to generate proper answers:</p>The context only contains information about CVE-2019-11756, which is...
<p>The observed behavior prompted us with an interesting question:</p>How could we use the fact that users often include close-to-exact CVE codes in their queries to enhance the accuracy of our search results?<p>To solve this, we approached the issue as a search challenge. We hypothesized that by emphasizing the <code>title</code> field matching for such articles, which directly contain the CVE codes, we could significantly improve the precision of our search results. This led to a strategic decision to conditionally boost the weighting of title matches in our search algorithm. By implementing this focused adjustment, we refined our query strategy as follows:</p>    ...
    should: [
        // Additional boosting for CVEs.
        {
          bool: {
            filter: {
              term: {
                category: 'cve',
              },
            },
            must: {
              match: {
                title: {
                  query: queryText,
                  boost: 10,
                },
              },
            },
          },
        },
        // BM-25 based search.
        {
          multi_match: {
             ...
<p>As a result, we experienced much better hits for CVE-related use cases, ensuring that <code>CVE-2016-1837</code>, <code>CVE-2019-11756</code> and <code>CVE-2014-6439</code> are top 3:</p>{
  ...
  "hits": [
    {
      "_index": "search-knowledge-articles",
      "_id": "...",
      "_score": 181.63962,
      "_source": {
        "title": "CVE-2019-11756",
        "summary": "...",
        "body": "...",
        "category": "cve"
      }
    },
    {
      "_index": "search-knowledge-articles",
      "_id": "...",
      "_score": 175.13728,
      "_source": {
        "title": "CVE-2014-6439",
        "summary": "...",
        "body": "...",
        "category": "cve"
      }
    },
    {
      "_index": "search-knowledge-articles",
      "_id": "...",
      "_score": 152.9553,
      "_source": {
        "title": "CVE-2016-1837",
        "summary": "...",
        "body": "...",
        "category": "cve"
      }
    }
  ]
}
<p>And thus generating a much better response by the LLM:</p>🤖 The implications of the CVEs mentioned are as follows: (...)
<p>Lovely! By tuning our Hybrid Search approach, we significantly improved our performance with a pretty simple, but mostly effective <em>Bob's Your Uncle</em> solution (like some folks would say)! This improvement underscores that while semantic search is a powerful tool, understanding and leveraging user intent is crucial for optimizing search results and overall chat experience in your business reality. With that in mind, let's dive into the next challenge!</p><h3>#2: Product versions</h3><p>As we delved deeper into the challenges, another significant issue emerged with queries related to specific versions. Users frequently inquire about features, migration guides, or version comparisons, but our initial search responses were not meeting expectations. For instance, let's take the following question:</p>👨🏽 Can you compare Elasticsearch versions 8.14.3 and 8.14.2?
<p>Our initial query approach would return the following top 3:</p><ul><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/hadoop/current/eshadoop-8.14.1.html">Elasticsearch for Apache Hadoop version 8.14.1 | Elasticsearch for Apache Hadoop [8.14] | Elastic</a>;</p></li><li><p><a href="https://www.elastic.co/guide/en/observability/current/apm-release-notes-8.14.html">APM version 8.14 | Elastic Observability [8.14] | Elastic</a><em>;</em></p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/hadoop/8.14/eshadoop-8.14.3.html">Elasticsearch for Apache Hadoop version 8.14.3 | Elasticsearch for Apache Hadoop [8.14] | Elastic</a>.</p></li></ul><p>Corresponding to the following <code>_search</code> response:</p>{
  ...
  "hits": [
    {
      "_index": "search-docs",
      "_id": "6807c4cf67ad0a52e02c4c2ef436194d2796faa454640ec64cc2bb999fe6633a",
      "_score": 29.79520,
      "_source": {
        "title": "Elasticsearch for Apache Hadoop version 8.14.1 | Elasticsearch for Apache Hadoop [8.14] | Elastic",
        "summary": "...",
        "body": "...",
        "url": "https://www.elastic.co/guide/en/elasticsearch/hadoop/current/eshadoop-8.14.1.html",
        "category": "documentation"
      }
    },
    {
      "_index": "search-docs",
      "_id": "9233c6b53a9d2e8b5daf1f68e7d82590516f0dfc23d979783ea1146854dc9d63",
      "_score": 27.19520,
      "_source": {
        "title": "APM version 8.14 | Elastic Observability [8.14] | Elastic",
        "summary": "...",
        "body": "...",
        "url": "https://www.elastic.co/guide/en/observability/current/apm-release-notes-8.14.html",
        "category": "documentation"
      }
    },
    {
      "_index": "search-docs",
      "_id": "392a39dd81933523d6018cac5ae4fc38b31de4ae3e1c0f1655d98191133c4e8a",
      "_score": 26.7618,
      "_source": {
        "title": "Elasticsearch for Apache Hadoop version 8.14.3 | Elasticsearch for Apache Hadoop [8.14] | Elastic",
        "summary": "...",
        "body": "...",
        "url": "https://www.elastic.co/guide/en/elasticsearch/hadoop/8.14/eshadoop-8.14.3.html",
        "category": "documentation"
      }
    }
  ]
}
<p>Being irrevocably irrelevant, they ended up resulting in a completely uninformed answer from the chatbot, affecting the overall user experience and trust in the Support AI Assistant:</p>🤖 As of my last update, Elasticsearch 8.14.2 and 8.14.3 had not been released, and there were no official release notes available for these versions. Elasticsearch follows a rapid release cycle, and new versions often include bug fixes, performance improvements, and occasionally new features. However, without specific release notes or announcements from Elastic, I cannot provide details on what is new in these hypothetical versions.

To find the most up-to-date information on Elasticsearch releases, you should check the official Elastic website or their GitHub repository for the latest release notes. Release notes are typically published alongside the software release and provide a comprehensive list of changes, including new features, enhancements, bug fixes, and known issues
<p>Further investigating the issue we collected valuable insights. By replaying the query and looking into the search results, we noticed three serious problems with our crawled <em>Product Documentation</em> data that were contributing to the overall bad performance:</p><ol><li><p><a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#a-inaccurate-semantic-matching"><strong>Inaccurate semantic matching</strong></a>: Semantically, we definitely missed the shot. Why would we match against such specific articles, including two specifically about Apache Hadoop, when the question was so much broader than Hadoop?</p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#b-multiple-versions-same-articles"><strong>Multiple versions, same articles</strong></a>: Going further down on the hits of the initially asked question, we often noticed multiple versions for the same articles, with close to exactly the same content. That often led to a top 3 cluttered with irrelevant matches!</p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#c-wrong-versions-being-returned"><strong>Wrong versions being returned</strong></a>: It's fair to expect that having both <em>8.14.1</em> and <em>8.14.2</em> versions of the <em>Elasticsearch for Apache Hadoop</em> article, we'd return the latter for our query – but that just wasn't happening consistently.</p></li></ol><p>From the impact perspective, we had to stop and solve those – else, a considerable part of user queries would be affected. Let's dive into the approaches taken to solve both!</p><h4>A. Inaccurate semantic matching</h4><p>After some examination into our data, we've discovered that the root of our semantic matching issue lived in the fact that the <code>summary</code> field for <em>Product Documentation-type</em> articles generated upon ingestion by the crawler was just the first few characters of the <code>body</code>. This redundancy misled our semantic model, causing it to generate vector embeddings that did not accurately represent the document's content in relation to user queries.</p><p>As a data problem, we had to solve this problem in the data domain: by leveraging the use of GenAI and the GPT4 model, we made a team decision to craft a new AI Enrichment Service – introduced in the <a href="https://search-labs.elastic.co/search-labs/blog/genai-customer-support-building-a-knowledge-library#enriching-document-sources">2nd installment of this blog series</a>. We decided to create our own tool for a few specific reasons:</p><ul><li><p>We had unused PTU resources available. Why not use them?</p></li><li><p>We needed this data gap filled quickly, as this was probably the greatest relevance detractor.</p></li><li><p>We wanted a fully customizable approach to make our own experiments.</p></li></ul><p>Modeled to be generic, our usage for it boils down to generating four new fields for our data into a new index, using <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/ingest-enriching-data.html"><em>Enrich Processors</em></a> to make them available to the respective documents on the target indices upon ingestion. Here's a quick view into the specification for each field to be generated:</p>const fields: FieldToGenerate[] = [
  {
    // A one-liner summary for the article.
    name: 'ai_subtitle', 
    strategy: GenerationStrategy.AbstractiveSummarizer,
  },
  {
    // A longer summary for the article.
    name: 'ai_summary', 
    strategy: GenerationStrategy.AbstractiveSummarizer,
  },
  {
    // A list of questions answered by the article.
    name: 'ai_questions_answered', 
    strategy: GenerationStrategy.QuestionSummarizer,
  },
  {
    // A condensed list of tags for the article.
    name: 'ai_tags',
    strategy: GenerationStrategy.TagsSummarizer,
  }
];
<p>After generating those fields and setting up the index <em>Enrich Processors</em>, the underlying RAG-search indices were enriched with a new <code>ai_fields</code> object, also making ELSER embeddings available under <code>ai_fields.ml.inference</code>:</p>{
  ...
  "_source": {
    "product_name": "Elasticsearch",
    "version": "8.14",
    "url": "https://www.elastic.co/guide/en/elasticsearch/hadoop/8.14/eshadoop-8.14.1.html",
    "ai_fields": {
      "ai_summary": "ES-Hadoop 8.14.1; tested against Elasticsearch 8.14.1. ES-Hadoop 8.14.1 is a compatibility release, aligning with Elasticsearch 8.14.1. This version ensures seamless integration and operation with Elasticsearch's corresponding version, maintaining feature parity and stability across the Elastic ecosystem.",
      "ai_subtitle": "ES-Hadoop 8.14.1 Compatibility Release",
      "ai_tags": [
        "Elasticsearch",
        "ES-Hadoop",
        "Compatibility",
        "Integration",
        "Version 8.14.1"
      ],
      "source_id": "6807c4cf67ad0a52e02c4c2ef436194d2796faa454640ec64cc2bb999fe6633a",
      "ai_questions_answered": [
        "What is ES-Hadoop 8.14.1?",
        "Which Elasticsearch version is ES-Hadoop 8.14.1 tested against?",
        "What is the purpose of the ES-Hadoop 8.14.1 release?"
      ],
      "ml": {
        "inference": {
          "ai_subtitle_expanded": {...},
          "ai_summary_expanded": {...},
          "ai_questions_answered_expanded": {...}
        }
      }
    }
  }
  ...
}
<p>Now, we can tune the query to use those fields, making for better overall semantic and keyword matching:</p>   ...
   // BM-25 Component. 
   {
      multi_match: {
        ...
        type: 'cross_fields',
        fields: [
          ...
          // Adding the `ai_fields` to the `cross_fields` matcher.
          'ai_fields.ai_subtitle',
          'ai_fields.ai_summary',
          'ai_fields.ai_questions_answered',
          'ai_fields.ai_tags',
        ],
      },
   },
   {
      multi_match: {
        ...
        type: 'phrase',
        fields: [
          ...
          // Adding the `ai_fields` to the `phrase` matcher.
          'ai_fields.ai_subtitle.stem',
          'ai_fields.ai_summary.stem',
          'ai_fields.ai_questions_answered.stem',
        ],
      },
   },
   ...
   // Semantic Search Component.
   {
      text_expansion: {
        // Adding `text_expansion` queries for `ai_fields` embeddings.
        'ai_fields.ml.inference.ai_subtitle_expanded.predicted_value': {
          model_id: '.elser_model_2',
          model_text: queryText,
        },
      },
    },
    {
      text_expansion: {
        'ai_fields.ml.inference.ai_summary_expanded.predicted_value': {
          model_id: '.elser_model_2',
          model_text: queryText,
        },
      },
    },
    {
      text_expansion: {
        'ai_fields.ml.inference.ai_questions_answered_expanded.predicted_value':
          {
            model_id: '.elser_model_2',
            model_text: queryText,
          },
      },
    },
    ...
<p>Single-handedly, that made us much more relevant. More than that – it also opened a lot of new possibilities to use the AI-generated data throughout our applications – matters of which we'll talk about in future blog posts.</p><p>Now, before retrying the query to check the results: <strong>what about the multiple versions problem?</strong></p><h4>B. Multiple versions, same articles</h4><p>When duplicate content infiltrates these top positions, it diminishes the value of the data pool, thereby diluting the effectiveness of GenAI responses and leading to a suboptimal user experience. In this context, a significant challenge we encountered was <strong>the presence of multiple versions of the same article.</strong> This redundancy, while contributing to a rich collection of version-specific data, often cluttered the essential data feed to our LLM, reducing the diversity of it and therefore undermining the response quality.</p><p>To address the problem, we employed the</p><p><em>Elasticsearch API</em> <code>collapse</code> parameter, sifting through the noise and prioritizing only the most relevant version of a single content. To do that, we computed a new <code>slug</code> field into our <em>Product Documentation</em> crawled documents to identify different versions of the same article, using it as the <em>collapse field</em> (or <em>key</em>).</p><p></p><p>Taking the <em>Sort search results</em> <a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.14/sort-search-results.html">documentation page</a> as an example, we have two versions of this article being crawled:</p><ul><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.14/sort-search-results.html">Sort search results | Elasticsearch Guide [8.14] | Elastic</a></p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/7.17/sort-search-results.html">Sort search results | Elasticsearch Guide [7.17] | Elastic</a></p></li></ul><p>Those two will generate the following <code>slug</code>:</p>guide-en-elasticsearch-reference-sort-search-results<p>Taking advantage of that, we can now tune the query to use <code>collapse</code>:</p>...
const searchQuery = {
  index: "knowledge-search",
  body: {
    ...
    query: {...},
    collapse: {
      // This is a "field alias" that will point to the `slug` field for product docs.
      field: "collapse_field" 
    }
  }
};
...
<p>As a result, we'll now only show the top-scored documentation in the search results, which will definitely contribute to increasing the diversity of knowledge being sent to the LLM.</p><h4>C. Wrong versions being returned</h4><p>Similar to the <a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#1-cves-common-vulnerabilities-and-exposures">CVE matching problem</a>, we can boost results based on the specific versions being mentioned, allied with the fact that <code>version</code> is a separate field in our index. To do that, we used the following simple regex-based function to pull off versions directly from the user question:</p>/**
 * Extracts versions from the query text.
 * @param queryText The user query (or question).
 * @returns Array of versions found in the query text.
 * @example getVersionsFromQueryText("What's new in 8.14.2? And 8.14.3?") =&gt; ['8.14.2', '8.14.3'] 
 */
const getVersionsFromQueryText = (queryText: string) : string[] =&gt; {
  let results = null;
  const versions = [];
  // Will match `x.y.z` or `x.y` 
  const versionRegex = /\b\d+\.\d+(\.\d+)?\b/gi;

  while ((results = versionRegex.exec(queryText))) {
    const [version] = results;
    versions.push(version);
  }

  return versions;
}; 
<p>We then add one more query to the <code>should</code> clause, boosting the <code>version</code> field accordingly and getting the right versions to the top (whenever they're mentioned):</p>    // Additional boosting for versions.
    ...(versions.length &gt; 0 ? [
    {
      bool: {
        filter: {
          term: {
            category: 'documentation',
          },
        },
        must: {
          terms: {
            version: versions,
            boost: 4,
          },
        },
      },
   }] : []),
    // BM-25 Component.
   {
      multi_match: {
        ...
        type: 'cross_fields',
        fields: [
          ...
    ...
<p>With <a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#a-inaccurate-semantic-matching"><strong>A</strong></a>, <a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#b-multiple-versions-same-articles"><strong>B</strong></a> and <a href="https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance#c-wrong-versions-being-returned"><strong>C</strong></a> solved, we're probably ready to see some strong results!</p><h3>Let's replay the question!</h3><p>By replaying the previously tried question:</p>👨🏽 Can you compare Elasticsearch versions 8.14.3 and 8.14.2?
<p>And therefore running the Elasticsearch query once again, we get dramatically better results consisting of the following articles:</p><ul><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/master/release-notes-8.14.3.html">Elasticsearch version 8.14.3 | Elasticsearch Guide [master] | Elastic</a></p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/master/release-notes-8.14.2.html">Elasticsearch version 8.14.2 | Elasticsearch Guide [master] | Elastic</a></p></li><li><p><a href="https://www.elastic.co/guide/en/elasticsearch/reference/8.14/es-release-notes.html">Release notes | Elasticsearch Guide [8.14] | Elastic</a></p></li></ul><p>Consequently, we have a better answer generated by the LLM. More powerful than that – in the context of this conversation, the LLM is now conscious about versions of Elasticsearch that are newer than the model's cut-off date, crafting correct answers around those:</p>🤖 Certainly! Here's a formatted summary of the changes in Elasticsearch versions 8.14.2 and 8.14.3:

**Elasticsearch 8.14.2:**

- **Known Issue:**
  - Upgrading clusters from version 8.11.4 or earlier with non-master-eligible nodes may result in new functionality not being registered. This is fixed in version 8.15.0.

- **Bug Fixes:**
  - **Data streams:** Fixed lazy rollover request to rollover the target data stream once.
  - **ES|QL:** Addressed ESQL cancellation for exchange requests and issues with SingleValueQuery.LuceneQuery.
  - **Infra/Scripting:** Added array size check before returning array item in script doc values.
  - **Infra/Settings:** Ensured file settings readiness is guarded on file settings support.
  - **Machine Learning:** Resolved IndexOutOfBoundsException during inference.
  - **Mapping:** Re-defined index.mapper.dynamic setting for smoother 7.x to 8.x upgrades.
  - **Ranking:** Corrected the 'from' parameter when using sub_searches and rank.
  - **Search:** Supported hexstring for byte painless scorers and fixed tracking of collapse with docvalue_fields.

**Elasticsearch 8.14.3:**

- **Known Issue:**
  - Same as in 8.14.2 regarding the upgrade process and non-master-eligible nodes.

- **Bug Fixes:**
  - **Cluster Coordination:** Ensured tasks preserve versions in MasterService.
  - **ES|QL:** Implemented a compute listener.
  - **Mapping:** Set ignore_malformed only for the @timestamp automatically.
  - **TSDB:** Prohibited setting or updating the index.time_series.end_time in normal indices.

For both versions, users should be aware of the known issue related to upgrading clusters and follow the recommended resolution until the fix is available in version 8.15.0.

**URL References:**
- [Elasticsearch 8.14.2 Release Notes](https://www.elastic.co/guide/en/elasticsearch/reference/master/release-notes-8.14.2.html)
- [Elasticsearch 8.14.3 Release Notes](https://www.elastic.co/guide/en/elasticsearch/reference/master/release-notes-8.14.3.html)
<p>Exciting, right? But how can we quantify the improvements in our query at this point? <strong>Let's see the numbers together!</strong></p><h2>Measuring success</h2><p>To assess the performance implied by our changes, we've compiled a test suite based on user behavior, each containing a question plus a curated list of results that are considered relevant to answer it. Those will cover a wide wide range of subjects and query styles, reflecting the diverse needs of our users. Here's a complete look into it:</p>...
const initialCases: Array&lt;TestCase&gt; = [
  {
    query: 'Can you compare Elasticsearch versions 8.14.3 and 8.14.2?',
    expectedResults: [...], // Elasticsearch version 8.14.3 | Elasticsearch Guide | Elastic, Elasticsearch version 8.14.2 | Elasticsearch Guide | Elastic.
  },
  {
    query: "What are the implications of CVE's 2019-10202, 2019-11756, 2019-15903?",
    expectedResults: [...], // CVE-2016-1837; CVE-2019-11756; CVE-2014-6439. 
  },
  {
    query: 'How to run the support diagnostics tool?',
    expectedResults: [...], // How to install and run the support diagnostics troubleshooting utility; How to install and run the ECK support diagnostics utility.
  },
  {
    query: 'How can I create data views in Kibana via API?',
    expectedResults: [...], // Create data view API | Kibana Guide | Elastic; How to create Kibana data view using api; Data views API | Kibana Guide | Elastic.
  },
  {
    query: 'What would the repercussions be of deleting a searchable snapshot and how would you be able to recover that index?',
    expectedResults: [...], // The repercussions of deleting a snapshot used by searchable snapshots; Does delete backing index delete the corresponding searchable snapshots, and vice versa?; Can one use a regular snapshot to restore searchable snapshot indices?; [ESS] Can deleted index data be recovered Elastic Cloud / Elasticsearch Service?.
  },
  {
    query: 'How can I create a data view in Kibana?',
    expectedResults: [...], // Create a data view | Kibana Guide | Elastic; Create data view API | Kibana Guide [8.2] | Elastic; How to create Kibana data view using api.
  },
  {
    query: 'Do we have an air gapped version of the Elastic Maps Service?',
    expectedResults: [...], // Installing in an air-gapped environment | Elastic Installation and Upgrade Guide [master] | Elastic; Connect to Elastic Maps Service | Kibana Guide | Elastic; 1.6.0 release highlights | Elastic Cloud on Kubernetes | Elastic.
  },
  {
    query: 'How to setup an enrich processor?',
    expectedResults: [...], // Set up an enrich processor | Elasticsearch Guide | Elastic; Enrich processor | Elasticsearch Guide | Elastic; Enrich your data | Elasticsearch Guide | Elastic.
  },
  {
    query: 'How to use index lifecycle management (ILM)?',
    expectedResults: [...], // Tutorial: Automate rollover with ILM | Elasticsearch Guide | Elastic; ILM: Manage the index lifecycle | Elasticsearch Guide | Elastic; ILM overview | Elasticsearch Guide | Elastic.
  },
  {
    query: 'How to rotate my ECE UI proxy certificates?',
    expectedResults: [...], // Manage security certificates | Elastic Cloud Enterprise Reference | Elastic; Generate ECE Self Signed Proxy Certificate; ECE Certificate Rotation (2.6 -&gt; 2.10).
  },
  {
    query:
      'How to rotate my ECE UI proxy certificates between versions 2.6 and 2.10?',
    expectedResults: [...], // ECE Certificate Rotation (2.6 -&gt; 2.10); Manage security certificates | Elastic Cloud Enterprise Reference | Elastic; Generate ECE Self Signed Proxy Certificate.
  }
];
...
<p>But how do we turn those test cases into quantifiable success? To this end, we have employed Elasticsearch's <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-rank-eval.html">Ranking Evaluation API</a> alongside with the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/search-rank-eval.html#k-precision">Precision at K (P@K)</a> metric to determine how many relevant results are returned between the first <em>K</em> hits of a query. As we're interested in the top 3 results being fed into the LLM, we're making K = 3 here.</p><p>To automate the computation of this metric against our curated list of questions and effectively assess our performance gains, we used <em>TypeScript/Node.js</em> to create a simple script wrapping everything up. First, we define a function to make the corresponding <em>Ranking Evaluation</em> API calls:</p>const rankingEvaluation = async (
    // The query to execute ("before" or "after").
    getSearchRequestFn: (queryText: string) =&gt; string
) =&gt;
    const testSuite = getTestSuite();
    const rankEvalResult = await elasticsearchClient.rankEval({
      index: 'knowledge-search',
      body: {
    metric: {
      precision: {
        k: 3,
        relevant_rating_threshold: 1,
      },
    },
    // For each test case, we'll have one item here.
    requests: testSuite.map((testCase) =&gt; ({
      id: testCase.queryText,
      request: getSearchRequestFn(testCase.queryText),
      ratings: testCase.expectedResults.map(({ _id, _index }) =&gt; ({
        _index,
        _id,
        rating: 1, // A value &gt;= 1 means relevant.
      })),
    })),
      },
    });
    // Return a normalized version of the data.
    return transformRankEvalResult(rankEvalResult);
}
<p>After that, we need to define the search queries <em>before</em> and <em>after</em> the optimizations:</p>// Before the optimizations.
const getSearchRequestBefore = (queryText: string): any =&gt; ({
  query: {
    bool: {
      should: [
        {
          multi_match: {
            query: queryText,
            minimum_should_match: '1&lt;-1 3&lt;49%',
            type: 'cross_fields',
            fields: ['title', 'summary', 'body', 'id'],
          },
        },
        {
          multi_match: {
            query: queryText,
            type: 'phrase',
            boost: 9,
            fields: [
              'title.stem',
              'summary.stem',
              'body.stem',
            ],
          },
        },
        {
          text_expansion: {
            'ml.inference.title_expanded.predicted_value': {
              model_id: '.elser_model_2',
              model_text: queryText,
            },
          },
        },
        {
          text_expansion: {
            'ml.inference.summary_expanded.predicted_value': {
              model_id: '.elser_model_2',
              model_text: queryText,
            },
          },
        },
      ],
    },
  },
});

// After the optimizations.
const getSearchRequestAfter = (queryText: string): any =&gt; {
  const versions = getVersionsFromQueryText(queryText);
  const matchesKeywords = [
    {
      multi_match: {
        query: queryText,
        minimum_should_match: '1&lt;-1 3&lt;49%',
        type: 'cross_fields',
        fields: [
          'title',
          'summary',
          'body',
          'id',
          'ai_fields.ai_subtitle',
          'ai_fields.ai_summary',
          'ai_fields.ai_questions_answered',
          'ai_fields.ai_tags',
        ],
      },
    },
    {
      multi_match: {
        query: queryText,
        type: 'phrase',
        boost: 9,
        slop: 0,
        fields: [
          'title.stem',
          'summary.stem',
          'body.stem',
          'ai_fields.ai_subtitle.stem',
          'ai_fields.ai_summary.stem',
          'ai_fields.ai_questions_answered.stem',
        ],
      },
    },
  ];

  const matchesSemantics = [
    {
      text_expansion: {
        'ml.inference.title_expanded.predicted_value': {
          model_id: '.elser_model_2',
          model_text: queryText,
        },
      },
    },
    {
      text_expansion: {
        'ml.inference.summary_expanded.predicted_value': {
          model_id: '.elser_model_2',
          model_text: queryText,
        },
      },
    },
    {
      text_expansion: {
        'ai_fields.ml.inference.ai_subtitle_expanded.predicted_value': {
          model_id: '.elser_model_2',
          model_text: queryText,
        },
      },
    },
    {
      text_expansion: {
        'ai_fields.ml.inference.ai_summary_expanded.predicted_value': {
          model_id: '.elser_model_2',
          model_text: queryText,
        },
      },
    },
    {
      text_expansion: {
        'ai_fields.ml.inference.ai_questions_answered_expanded.predicted_value':
          {
            model_id: '.elser_model_2',
            model_text: queryText,
          },
      },
    },
  ];

  const matchesCvesAndVersions = [
    {
      bool: {
        filter: {
          term: {
            category: 'cve',
          },
        },
        must: {
          match: {
            title: {
              query: queryText,
              boost: 10,
            },
          },
        },
      },
    },
    ...(versions.length &gt; 0
      ? [
          {
            bool: {
              filter: {
                term: {
                  category: 'documentation',
                },
              },
              must: {
                terms: {
                  version: versions,
                  boost: 4,
                },
              },
            },
          },
        ]
      : []),
  ];

  return {
    query: {
      bool: {
        should: [
          ...matchesKeywords,
          ...matchesSemantics,
          ...matchesCvesAndVersions,
        ]
      },
    },
    collapse: {
      // Alias to the collapse key for each underlying index. 
      field: 'collapse_field' 
    },
  };
};
<p>Then, we'll output the resulting metrics for each query:</p>const [rankEvaluationBefore, rankEvaluationAfter] =
  await Promise.all([
    rankingEvaluation(getSearchRequestBefore), // The "before" query.
    rankingEvaluation(getSearchRequestAfter), // The "after" query.
  ]);

console.log(`Before -&gt; Precision at K = 3 (P@K):`);
console.table(rankEvaluationBefore);

console.log(`After -&gt; Precision at K = 3(P@k):`);
console.table(rankEvaluationAfter);

// Computing the change in P@K.
const metricScoreBefore = rankEvaluationBefore.getMetricScore();
const metricScoreAfter = rankEvaluationAfter.getMetricScore();

const percentDifference =
  ((metricScoreAfter - metricScoreBefore) * 100) / metricScoreBefore;

console.log(`Change in P@K: ${percentDifference.toFixed(2)}%`);
<p>Finally, by running the script against our <em>development</em> Elasticsearch instance, we can see the following output demonstrating the P@K or (P@3) values for each query, <em>before</em> and <em>after</em> the changes. That is – how many results on the top 3 are considered relevant to the response:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta011b61ca9ef32e9/6a17d7be7b54f914718b371c/c89071c3194eb69f251b6716ea6341a6e423c684-1061x580.png" alt="Script output for the ranking evaluation using the P@K metric" /><h2>Improvements observed</h2><p>As an archer carefully adjusts for a precise shot, our recent efforts into relevance have brought considerable improvements in precision over time. Each one of the previous enhancements, in sequence, were small steps towards achieving better accuracy in our RAG-search results, and overall user experience. Here's a look at how our efforts have improved performance across various queries:</p><h4>Before and after – <code>P@K</code></h4>Relevant results in the top 3: <code>❌ = 0</code>, <code>🥉 = 1</code>, <code>🥈 = 2</code>, <code>🥇 = 3</code>.<p>Query Description</p><p>P@K Before</p><p>P@K After</p><p>Change</p><p>Support Diagnostics Tool</p><p>0.333 🥉</p><p>1.000 🥇</p><p>+200%</p><p>Air Gapped Maps Service</p><p>0.333 🥉</p><p>0.667 🥈</p><p>+100%</p><p>CVE Implications</p><p>0.000 ❌</p><p>1.000 🥇</p><p>∞</p><p>Enrich Processor Setup</p><p>0.667 🥈</p><p>0.667 🥈</p><p>0%</p><p>Proxy Certificates Rotation</p><p>0.333 🥉</p><p>0.333 🥉</p><p>0%</p><p>Proxy Certificates Version-specific Rotation</p><p>0.333 🥉</p><p>0.333 🥉</p><p>0%</p><p>Searchable Snapshot Deletion</p><p>0.667 🥈</p><p>1.000 🥇</p><p>+50%</p><p>Index Lifecycle Management Usage</p><p>0.667 🥈</p><p>0.667 🥈</p><p>0%</p><p>Creating Data Views via API in Kibana</p><p>0.333 🥉</p><p>0.667 🥈</p><p>+100%</p><p>Kibana Data View Creation</p><p>1.000 🥇</p><p>1.000 🥇</p><p>0%</p><p>Comparing Elasticsearch Versions</p><p>0.000 ❌</p><p>0.667 🥈</p><p>∞</p><p>Maximum Bucket Size in Aggregations</p><p>0.000 ❌</p><p>0.333 🥉</p><p>∞</p><p><strong>Average </strong><strong><code>P@K</code></strong><strong> Improvement: +78.41% 🏆🎉</strong>. Let's summarize a few observations about our results:</p><p><strong>Significant Improvements</strong>: With the measured overall <strong>+78.41%</strong> of relevance increase, the following queries – <em>Support Diagnostics Tool</em>, <em>CVE implications</em>, <em>Searchable Snapshot Deletion, Comparing Elasticsearch Versions</em> – showed substantial enhancements. These areas not only reached the <em>podium</em> of search relevance but did so with flying colors, significantly outpacing their initial performances!</p><p><strong>Opportunities for Optimization</strong>: Certain queries like the <em>Enrich Processor Setup</em>, <em>Kibana Data View Creation</em> and <em>Proxy Certificates Rotation</em> have shown reliable performances, without regressions. These results underscore the effectiveness of our core search strategies. However, those remind us that precision in search is an ongoing effort. These static results highlight where we'll focus our efforts to sharpen our aim throughout the next iterations. As we continue, we'll also expand our test suite, incorporating more diverse and meticulously selected use cases to ensure our enhancements are both relevant and robust.</p><h2>What's next? 🔎</h2><p>The path ahead is marked by opportunities for further gains, and with each iteration, we aim to push the RAG implementation performance and overall experience even higher. With that, let's discuss areas that we're currently interested in!</p><ol><li><p><strong>Our data can be futher optimized for search</strong>: Although we have a large base of sources, we observed that having semantically close search candidates often led to less effective chatbot responses. Some of the crawled pages aren't really valuable, and often generate noise that impacts relevance negatively. To solve that, we can curate and enhance our existing knowledge base by applying a plethora of techniques, making it lean and effective to ensure an optimal search experience.</p></li><li><p><strong>Chatbots must handle conversations – and so must RAG searches</strong>: It's common user behavior to ask follow-up questions to the chatbot. A question asking "How to configure Elasticsearch on a Linux machine?" followed by "What about Windows?" should query something like "How to configure Elasticsearch on a Linux machine?" (not the raw 2nd question). The RAG query approach should find the most relevant content regarding the entire context of the conversation.</p></li><li><p><strong>Conditional context inclusion</strong>: By extracting the semantic meaning of the user question, it would be possible to conditionally include pieces of data as context, saving token limits, making the generated content even more relevant, and potentially saving <em>round trips</em> for search and external services.</p></li></ol><h2>Conclusion</h2><p>In this installment of our series on GenAI for Customer Support, we have thoroughly explored the enhancements to the Retrieval-Augmented Generation (RAG) search within Elastic's customer support systems. By refining the interaction between large language models and our search algorithms, we have successfully elevated the precision and effectiveness of the Support AI Assistant.</p><p>Looking ahead, we aim to further optimize our search capabilities and expand our understanding of user interactions. This continuous improvement will focus on refining our AI models and search algorithms to better serve user needs and enhance overall customer satisfaction.</p><p>Stay tuned for more insights and updates as we continue to push the boundaries of what's possible with AI in customer support, and don't forget to join us in our next discussion, where we'll explore how Observability plays a critical role in monitoring, diagnosing, and optimizing the performance and reliability of the Support AI Assistant as we scale!</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elser-rag-search-for-relevance</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Antonio Schönmann]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt975c900958571074/6a17d7c0abe0f277e2dfe85d/fd800d12d1c12abf68ccb7e8dd80ad7b62bec38c-1440x840.png" length="0" type="image/png"/>
    <pubDate>Thu, 22 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elastic Support Hub starts using semantic search]]></title>
    <description><![CDATA[We transitioned our Support Hub to semantic search, a more advanced search method that understands user intent rather than relying on keywords. This transition helps provide customers with more relevant search results in Elastic's support content.]]></description>
    <content:encoded><![CDATA[<p>We’re excited to share a recent enhancement made to the Elastic Support Hub: it’s now powered by semantic search!</p><p>But before we go into more detail on the changes we made to the Elastic® Support Hub and its impact on our customers, it's important that we take a moment to explain the concept of semantic search. At its core, semantic search is a method of search that uses AI to return more relevant search results. Take a look at this quick video explaining the concept:</p><p>As shown in the video, semantic search matches the <em>intent</em> of what the user searches to the content available rather than the <em>words</em>. You can read more about the AI behind it on our blog, <a href="https://www.elastic.co/search-labs/may-2023-launch-sparse-encoder-ai-model">Introducing Elastic Learned Sparse Encoder: Elastic’s AI model for semantic search</a>. The rest of this blog tells our story about moving the Elastic <a href="https://support.elastic.co/home">Support Hub</a> to semantic search.</p><h2>Why did we make this change?</h2><p>All technology news these days seems to have something to do with <a href="https://www.elastic.co/what-is/large-language-models">large language models</a> and <a href="https://www.elastic.co/what-is/generative-ai">generative AI</a>. Elastic is leading the charge with its <a href="https://www.elastic.co/elasticsearch/vector-database">vector database capabilities</a> and built-in natural language models. It makes sense that we should build our supporting applications on the same bleeding edge that our product lives on. By making this change now, we can provide feedback to our product development teams and make the product better for everyone.</p><h2>Biggest takeaway configuring semantic search</h2><p>As with most new technology innovations, it requires tearing down, replacing older code, and potentially updating underlying architecture. Our internal app development team faced these challenges head-on, and we are now in a much better position to iterate on any of Elasticsearch®’s new features. From our teams' point of view, there were two significant features that stood out in the setup process:</p><p>1. Considering ELSER, Elastic’s proprietary transformer model for semantic search, is a relatively new feature in Elasticsearch (8.8), our development team was happy to see a guided UI experience to enable Elasticsearch ingest pipelines with ELSER.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3ad9bab63a32532d/6a17d70d7b54f92cc48b3716/6d41f12865ca5fede4ce8cb41d6851b6ed69fa47-814x586.png" alt="" /><p>This allowed our developers to quickly add the necessary text expansion configuration to the ingestion pipeline that makes semantic search possible. This made the configuration experience much easier to get started and see results quicker.</p><p>2. A machine learning model like ELSER takes dedicated machine resources to run (minimum 4GB). Since we were already running on <a href="https://cloud.elastic.co/">Elastic Cloud</a>, we were able to enable dedicated machine learning (ML) nodes with autoscaling to accommodate our resource demands and see more consistent performance.</p><h2>Early evaluation of search results</h2><p>We are enabling various systems to help us to understand user queries, search results, and relevancy at scale. However, in our user testing, we can already see significant improvements in various queries. For example, we tested the phrase “How to index data into Elasticsearch” on both our standard full-text search and our new semantic search implementations.</p><p>Here is a side-by-side comparison of the two search methods.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt467faa649e4d2c87/6a17d70f1480093ce0b485b9/4249b79ed37ca62593e3a5466aeca92554d674c5-1440x692.png" alt="" /><p>While there isn’t a single article that explains all the ways you can index data (there are a lot), you can see how fundamentally different these results are. For full-text search, we have a mix of guides, troubleshooting articles, and a blog with matched keywords, but none of them answer the question of “how.” Or to say it differently, text search didn't capture the meaning (semantically) of the query and did its best to match keywords.</p><p>For the semantic search results, you can see blogs that generally relate to the indexing of data. What is even more interesting is the fourth returned result of “How to ingest data into Elasticsearch Service” as the term ingest is actually more relevant to the process of adding data to an index. Elastic’s out-of-the-box transformer model picked up on the semantic meaning of adding data to an index and returned more relevant results regardless of the exact keywords.</p><h2>What’s next?</h2><p>While we see this as a gigantic leap forward in our ability to provide customers with relevant search results, we know our work is not done. Over time, we will evaluate the data we have on terms searched, results, and articles read. This data will allow us to add <a href="https://www.elastic.co/guide/en/app-search/current/synonyms-guide.html">synonyms</a> and configure appropriate <a href="https://www.elastic.co/guide/en/app-search/current/relevance-tuning-guide.html">weights and boosts</a> to give you, our customers, the best experience when searching for Elastic content on <a href="https://support.elastic.co/home">support.elastic.co</a>.</p><p><a href="https://www.elastic.co/blog/elastic-knowledge-center-support-hub">&gt;&gt; Learn more about all the Support Hub has to offer.</a></p><p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p><p><em>In this blog post, we may have used or referred to third party generative AI tools, which are owned and operated by their respective owners. Elastic does not have any control over the third party tools and we have no responsibility or liability for their content, operation or use, nor for any loss or damage that may arise from your use of such tools. Please exercise caution when using AI tools with personal, sensitive or confidential information. Any data you submit may be used for AI training or other purposes. There is no guarantee that information you provide will be kept secure or confidential. You should familiarize yourself with the privacy practices and terms of use of any generative AI tools prior to use.</em></p><p><em>Elastic, Elasticsearch, ESRE, Elasticsearch Relevance Engine and associated marks are trademarks, logos or registered trademarks of Elasticsearch N.V. in the United States and other countries. All other company and product names are trademarks, logos or registered trademarks of their respective owners.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elastic-support-hub-uses-semantic-search</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elastic-support-hub-uses-semantic-search</guid>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Chris Blaisure]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbeb5dbab80ffb03d/6a17d7104202290f4229f449/7fce63922298d2a0fcc8a5f292db74d2e0fae0a1-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 16 Nov 2023 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>