Blog

The fastest work is the work you never do: 100x faster sorted queries in Elasticsearch

How Elasticsearch 9.6 lets ES|QL push its current TopN threshold into Lucene, and how that 100x shows up on some sorted queries and barely registers on others.

Get hands-on with Elasticsearch: Dive into our sample notebooks in the Elasticsearch Labs repo, start a free cloud trial, or try Elastic on your local machine now.

In Elasticsearch 9.6, a query filtering 346 million Nginx log documents dropped from around 500 seconds to 5 on a single node. That’s 100x faster, and the same query gains 37x on Elastic Cloud Serverless. The optimization is called min-competitive TopN.As Elasticsearch Query Language (ES|QL) accumulates result candidates, it raises the minimum value that Lucene has to return, so document ranges that can no longer improve the result are never loaded or decoded. You get the same 500 rows back, with hundreds of millions of documents left unread.

A sorted ES|QL query over 346 million log documents

FROM logs_*
| WHERE message LIKE "*request body too large*" OR message LIKE "*queue is full*"
| SORT @timestamp DESC
| LIMIT 500

Using real-world Nginx logs, the benchmark searched 346 million docs across 432 GB of total index data. It returned 500 results and took ~500 seconds. That's more than eight minutes.

What’s happening under the hood?

Before this optimization, candidate rows were loaded, decoded, and evaluated by the user filter before TopN could determine which of them could no longer contribute to the final result.

How min-competitive TopN prunes work in ES|QL

TopN learns something useful while the query is still running. Once its heap is full, the worst value currently in the heap becomes a competitive threshold: that is, any row that cannot beat that value cannot make the final result.

ES|QL can feed that threshold back into earlier stages of execution. Two changes make this work:

  1. Dynamic Lucene pushdown. As TopN finds better candidates, its competitive threshold improves. ES|QL pushes that threshold down to Lucene, allowing Lucene to skip candidates that cannot beat the current worst TopN value. This avoids loading, decoding, and evaluating rows that have no chance of making the result.

  2. Shared threshold across drivers. ES|QL processes data in parallel across multiple drivers. If each driver only uses its own local threshold, one driver cannot benefit from better candidates discovered by another. Sharing the best-known competitive threshold means a strong result found by one driver immediately helps all the others prune their remaining work.

Together, these changes move TopN from a late-stage filter to something that shapes query execution much earlier in the pipeline.

How does the min-competitive TopN optimization work?

The query asks for the top three rows matching a wildcard filter, sorted in descending order, across three segments of 10 rows each. Without the optimization, all 30 rows are loaded, decoded, and evaluated against the filter before TopN sees any of them, at a cost of 30 row loads and 15 TopN operations.

With dynamic pushdown, ES|QL maintains the TopN heap as it scans. The heap fills at the end of the first segment with [98, 51, 47], and its worst entry is pushed down to Lucene as the competitive threshold. That cuts the second segment to four rows out of 10, which improves the heap to [98, 83, 77] and tightens the threshold again, leaving just two rows to load in the third segment. This produces the same three results, with 16 row loads and 10 TopN operations.

Sharing the TopN threshold across drivers

One driver may find excellent recent candidates immediately. Without coordination, another driver doesn't know that and continues working with a weak threshold. Sharing the global worst competitive value lets that discovery benefit all drivers.

Benchmark results on a local node and Elastic Cloud Serverless

The min-competitive TopN optimization ships as generally available (GA) in Elasticsearch 9.6. For real-world Nginx data benchmarks (same query, same dataset, p50 of 51 measurements, with 20 warmup runs discarded):

Local node

Serverless (GCP)

Before optimization

~500s

~520s

Production implementation

~5s

~14s

Runtime improvement

~99%

~97%

Speedup

100x

37x

Benchmark hardware

Local node

Serverless (GCP)

CPU

Apple M4

Small GCP host

RAM

48 GB

8 GB

Java Virtual Machine (JVM) heap

24 GB

4 GB

Storage

Local external TB5 NVMe SSD

Cloud storage

Shards

100

100

When does this query optimization technique help most?

The biggest gains happen when TopN becomes competitive early. If good candidates appear at the start of execution, the threshold rises quickly and subsequent scans get pruned aggressively. Three conditions amplify this:

  1. N is small relative to the candidate population. Finding 500 rows out of 346 million documents gives the optimizer far more room to eliminate work than returning a large fraction of the matches does.

  2. The post-Lucene work is expensive enough that skipping a row saves something measurable, like avoiding loading large strings, decompression, or wildcard evaluation.

  3. Many candidates pass the coarse Lucene scan in the first place. If almost nothing matches the initial filter, there’s not much unnecessary work left to eliminate.

If matching rows are rare or if good candidates are discovered late, the threshold takes longer to become useful and the gain can be much smaller.

The current implementation applies to @timestamp sorted queries. Queries sorted by other fields don’t yet benefit from this optimization.

What's next for ES|QL query performance

The same idea has room to run in several directions, including:

  • Other numeric sort fields. The underlying approach isn’t inherently limited to timestamps, so extending it to other numeric sort fields could benefit a much broader range of sorted TopN queries.

  • Processing promising data first. How quickly the threshold becomes useful depends on how early TopN sees good candidates. Processing more promising data slices first could establish a strong threshold sooner, allowing more aggressive pruning of subsequent work.

  • Early termination. As the threshold strengthens during execution, it becomes possible to recognize when remaining ranges cannot produce a better result and to stop processing them entirely, rather than running through to completion.

  • Coordination across nodes. The current optimization shares competitive information across drivers within a node. Broader coordination across nodes or clusters could let strong candidates discovered in one part of a distributed query reduce work elsewhere.

  • Late materialization and result streaming. Competitive filtering and late materialization attack the same problem from different angles. Competitive filtering eliminates rows that cannot reach the final TopN as early as possible, while a fetch phase can defer loading expensive field values until the set of likely winners is much smaller. Result streaming could extend this further. If the engine can determine that some leading results can no longer be displaced, those rows could be returned before the full query completes. For interactive investigations, getting the first useful result fast matters as much as total query runtime.

What this means for Elasticsearch performance tuning

The min-competitive TopN optimization in Elasticsearch 9.6 shows how much work can be avoided when information discovered during query execution is fed back into earlier stages. In our Nginx log benchmark, that turned a roughly 500-second query into a roughly 5-second query on a single node.

This is also a starting point. Extending competitive filtering to more sort fields, finding strong candidates earlier, terminating work sooner, and eventually coordinating these decisions across nodes offer opportunities to push the same idea further.

The earlier we know what can still win, the less work we need to do for everything that cannot.

Frequently Asked Questions

How does min-competitive TopN make Elasticsearch queries faster?

Min-competitive TopN optimization is an ES|QL execution improvement in Elasticsearch 9.6 that feeds the growing result set back into earlier stages of query execution. As ES|QL accumulates candidates for a sorted LIMIT query, it dynamically raises the minimum value that Lucene needs to return, allowing it to skip document ranges that can no longer improve the final result. This reduces query time by up to 99% on large datasets.

Do all sorted LIMIT queries benefit equally from this optimization?

No. The improvement is specialized for @timestamp sorted TopN queries and depends on how quickly a useful competitive threshold is established, the requested TopN size, the distribution of matching rows, and the cost of loading and evaluating candidates.

Does this optimization change my query results?

No. The optimization changes how much work is needed to find the TopN, not which rows qualify for it. Candidates are skipped only when they cannot improve the current competitive result.

Does this work for Elasticsearch search API queries, too?

Not yet. The optimization described in this post applies to ES|QL sorted queries. A similar approach for the Elasticsearch search API is a known direction that the team is working on. The tracking issue #136267 covers the broader min-competitive optimization work, including potential future extensions.

How helpful was this content?

Related Content

The best LLM writes correct Elasticsearch ES|QL 59% of the time. Here's what breaks the other 41%.

The best LLM writes correct Elasticsearch ES|QL 59% of the time. Here's what breaks the other 41%.

Jeffrey Rengifo
Ask Elastic Agent Builder why it's slow: Natural-language trace analysis

Ask Elastic Agent Builder why it's slow: Natural-language trace analysis

Meghan Murphy
Columnar storage isn't a columnar database. What Columnar mode brings to Elasticsearch

Columnar storage isn't a columnar database. What Columnar mode brings to Elasticsearch

Yannis Roussos
Query rewrite rules in Elasticsearch: 2.3x faster wildcard scans

Query rewrite rules in Elasticsearch: 2.3x faster wildcard scans

Parker Timmins

Ready to build state of the art search experiences?

Sufficiently advanced search isn’t achieved with the efforts of one. Elasticsearch is powered by data scientists, ML ops, engineers, and many more who are just as passionate about search as you are. Let’s connect and work together to build the magical search experience that will get you the results you want.