<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Attila Sahi - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Attila Sahi - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/search-labs/author/attila-sahi</link>
    </image>
    <link>https://www.elastic.co/search-labs/author/attila-sahi</link>
    <atom:link href="https://www.elastic.co/search-labs/rss/author/attila-sahi.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Mon, 05 Oct 2026 19:36:59 GMT</lastBuildDate>
  <item>
    <title><![CDATA[The fastest work is the work you never do: 100x faster sorted queries in Elasticsearch]]></title>
    <description><![CDATA[How Elasticsearch 9.6 lets ES|QL push its current TopN threshold into Lucene, and how that 100x shows up on some sorted queries and barely registers on others.]]></description>
    <content:encoded><![CDATA[<p>In <a href="https://www.elastic.co/elasticsearch">Elasticsearch</a> 9.6, a query filtering 346 million Nginx log documents dropped from around 500 seconds to 5 on a single node. That’s 100x faster, and the same query gains 37x on <a href="https://www.elastic.co/docs/deploy-manage/deploy/elastic-cloud/serverless">Elastic Cloud Serverless</a>. The optimization is called <em>min-competitive TopN</em>.As <a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language (ES|QL)</a> accumulates result candidates, it raises the minimum value that Lucene has to return, so document ranges that can no longer improve the result are never loaded or decoded. You get the same 500 rows back, with hundreds of millions of documents left unread.</p><h2>A sorted ES|QL query over 346 million log documents</h2>FROM logs_*
| WHERE message LIKE "*request body too large*" OR message LIKE "*queue is full*"
| SORT @timestamp DESC
| LIMIT 500<p>Using real-world Nginx logs, the benchmark searched 346 million docs across 432 GB of total index data. It returned 500 results and took ~500 seconds. That's more than eight minutes.</p><p></p><h3>What’s happening under the hood?</h3><p></p><p>Before this optimization, candidate rows were loaded, decoded, and evaluated by the user filter before TopN could determine which of them could no longer contribute to the final result.
</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf091d4d4f50632fe/6abfaca2e7a20edb3234a965/esql-query-pipeline.png" alt="ES|QL query pipeline: Lucene scan, load and decode, filter, local TopN, final TopN, with the TopN decision last" /><p></p><h2>How min-competitive TopN prunes work in ES|QL</h2><p>TopN learns something useful while the query is still running. Once its heap is full, the worst value currently in the heap becomes a competitive threshold: that is, any row that cannot beat that value cannot make the final result.</p><p></p><p>ES|QL can feed that threshold back into earlier stages of execution. Two changes make this work:</p><p></p><ol><li><p><a href="https://www.elastic.co/search-labs/blog/esql-elasticsearch-8-19-9-1?#significant-performance-and-scalability-improvements"><strong>Dynamic Lucene pushdown</strong></a><strong>.</strong> As TopN finds better candidates, its competitive threshold improves. ES|QL pushes that threshold down to Lucene, allowing Lucene to skip candidates that cannot beat the current worst TopN value. This avoids loading, decoding, and evaluating rows that have no chance of making the result.</p></li><li><p><a href="https://github.com/elastic/elasticsearch/pull/142406"><strong>Shared threshold across drivers</strong></a><strong>.</strong> ES|QL processes data in parallel across multiple drivers. If each driver only uses its own local threshold, one driver cannot benefit from better candidates discovered by another. Sharing the best-known competitive threshold means a strong result found by one driver immediately helps all the others prune their remaining work.</p></li></ol><p></p><p>Together, these changes move TopN from a late-stage filter to something that shapes query execution much earlier in the pipeline.</p><h3>How does the min-competitive TopN optimization work?</h3><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7807646d99ccf5ff/6abfad1b0491ebddfc047acb/baseline-topn-query.png" alt="Baseline TopN query scanning all 30 rows across three Lucene segments to return the top 3 sorted results" /><p>The query asks for the top three rows matching a wildcard filter, sorted in descending order, across three segments of 10 rows each. Without the optimization, all 30 rows are loaded, decoded, and evaluated against the filter before TopN sees any of them, at a cost of 30 row loads and 15 TopN operations.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltca82b2e1bee9dfb7/6abfad6b23d193d2985b6ccb/dynamic-lucene-pushdown.png" alt="Dynamic Lucene pushdown cutting rows loaded from 30 to 16 as the TopN min-heap threshold tightens per segment" /><p>With dynamic pushdown, ES|QL maintains the TopN heap as it scans. The heap fills at the end of the first segment with [98, 51, 47], and its worst entry is pushed down to Lucene as the competitive threshold. That cuts the second segment to four rows out of 10, which improves the heap to [98, 83, 77] and tightens the threshold again, leaving just two rows to load in the third segment. This produces the same three results, with 16 row loads and 10 TopN operations.</p><h3>Sharing the TopN threshold across drivers</h3><p>One driver may find excellent recent candidates immediately. Without coordination, another driver doesn't know that and continues working with a weak threshold. Sharing the global worst competitive value lets that discovery benefit all drivers.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt71b29eda542cb155/6abfadc44e42ee2b277fce13/parallel-esql-drivers.png" alt="Parallel ES|QL drivers sharing a global TopN min-heap threshold to prune work across segments and shards" /><p></p><h2>Benchmark results on a local node and Elastic Cloud Serverless</h2><p>The min-competitive TopN optimization ships as generally available (GA) in Elasticsearch 9.6. For real-world Nginx data benchmarks (same query, same dataset, p50 of 51 measurements, with 20 warmup runs discarded):</p><p></p><p>
</p><p><strong>Local node</strong></p><p><strong>Serverless (GCP)</strong></p><p>Before optimization</p><p>~500s</p><p>~520s</p><p>Production implementation</p><p>~5s</p><p>~14s</p><p>Runtime improvement</p><p>~99%</p><p>~97%</p><p>Speedup</p><p>100x</p><p>37x</p><p></p><h3>Benchmark hardware</h3><p></p><p>
</p><p><strong>Local node</strong></p><p><strong>Serverless (GCP)</strong></p><p>CPU</p><p>Apple M4</p><p>Small GCP host</p><p>RAM</p><p>48 GB</p><p>8 GB</p><p>Java Virtual Machine (JVM) heap</p><p>24 GB</p><p>4 GB</p><p>Storage</p><p>Local external TB5 NVMe SSD</p><p>Cloud storage</p><p>Shards</p><p>100</p><p>100</p><h2>When does this query optimization technique help most?</h2><p>The biggest gains happen when TopN becomes competitive early. If good candidates appear at the start of execution, the threshold rises quickly and subsequent scans get pruned aggressively. Three conditions amplify this:</p><p></p><ol><li><p>N is small relative to the candidate population. Finding 500 rows out of 346 million documents gives the optimizer far more room to eliminate work than returning a large fraction of the matches does.</p></li><li><p>The post-Lucene work is expensive enough that skipping a row saves something measurable, like avoiding loading large strings, decompression, or wildcard evaluation.</p></li><li><p>Many candidates pass the coarse Lucene scan in the first place. If almost nothing matches the initial filter, there’s not much unnecessary work left to eliminate.</p></li></ol><p></p><p>If matching rows are rare or if good candidates are discovered late, the threshold takes longer to become useful and the gain can be much smaller.</p><p></p><p>The current implementation applies to <code>@timestamp</code> sorted queries. Queries sorted by other fields don’t yet benefit from this optimization.</p><h2>What's next for ES|QL query performance</h2><p>The same idea has room to run in several directions, including:</p><p></p><ul><li><p><strong>Other numeric sort fields. </strong>The underlying approach isn’t inherently limited to timestamps, so extending it to other numeric sort fields could benefit a much broader range of sorted TopN queries.</p></li></ul><p></p><ul><li><p><strong>Processing promising data first.</strong> How quickly the threshold becomes useful depends on how early TopN sees good candidates. Processing more promising data slices first could establish a strong threshold sooner, allowing more aggressive pruning of subsequent work.</p></li></ul><p></p><ul><li><p><strong>Early termination.</strong> As the threshold strengthens during execution, it becomes possible to recognize when remaining ranges cannot produce a better result and to stop processing them entirely, rather than running through to completion.</p></li></ul><p></p><ul><li><p><strong>Coordination across nodes.</strong> The current optimization shares competitive information across drivers within a node. Broader coordination across nodes or clusters could let strong candidates discovered in one part of a distributed query reduce work elsewhere.</p></li></ul><p></p><ul><li><p><strong>Late materialization and result streaming.</strong> Competitive filtering and late materialization attack the same problem from different angles. Competitive filtering eliminates rows that cannot reach the final TopN as early as possible, while a fetch phase can defer loading expensive field values until the set of likely winners is much smaller. Result streaming could extend this further. If the engine can determine that some leading results can no longer be displaced, those rows could be returned before the full query completes. For interactive investigations, getting the first useful result fast matters as much as total query runtime.</p></li></ul><p></p><h2>What this means for Elasticsearch performance tuning</h2><p>
The min-competitive TopN optimization in Elasticsearch 9.6 shows how much work can be avoided when information discovered during query execution is fed back into earlier stages. In our Nginx log benchmark, that turned a roughly 500-second query into a roughly 5-second query on a single node.</p><p></p><p>This is also a starting point. Extending competitive filtering to more sort fields, finding strong candidates earlier, terminating work sooner, and eventually coordinating these decisions across nodes offer opportunities to push the same idea further.</p><p></p><p>The earlier we know what can still win, the less work we need to do for everything that cannot.</p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-performance-tuning-esql-topn</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-performance-tuning-esql-topn</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Lucene]]></category>
    <dc:creator><![CDATA[Attila Sahi]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf092de2db9684039/6abfa9e5701a1a7aa18aa403/faster-query-runtime.png" length="0" type="image/png"/>
    <pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>