Blog

How Elasticsearch Serverless hollow shards cut indexing-node shutdowns by 30%

Idle indexing shards in Elasticsearch Serverless now drop their Lucene writers and segment readers from memory until the next write, freeing indexing-tier heap at the cost of a median 218ms wait on that first write.

Free yourself from operations with Elastic Cloud Serverless. Scale automatically, handle load spikes, and focus on building—start a 14-day free trial to test it out yourself!

You can follow these guides to build an AI-Powered search experience or search across business systems and software.

Elasticsearch Serverless now unloads idle indexing shards from memory. We call them hollow shards; the shard stays allocated, but its Lucene IndexWriter and segment readers are gone until the next write arrives. Hollow shards build on thin indexing shards, which already moved Lucene files off the local disk. In the first month in production, indexing-node shutdown times dropped by up to 30% at high percentiles across the fleet, and hundreds of thousands of shards went hollow. On one project with many idle data stream backing indices, shutdown time and indexing-tier heap improved by up to 3x. Search latency is unchanged, and the first write to a hollow shard waits a median 218ms while the shard reloads.

Thin shards write Lucene (the search library that Elasticsearch uses for indexing and search) segments locally and upload them to the object store as batched compound commits. Then they delete them from disk.

A thin shard that isn't ingesting and isn't expected to do so in the near future still holds an IndexWriter, segment readers, mappings, and the rest of its engine. At data stream scale, those idle engines make up most of the indexing heap and add time to the relocations that we need for balancing and for shutting nodes down during autoscaling.

A hollow indexing shard is a thin allocated primary that keeps only the last commit's metadata and doesn't create its IndexWriter or segment readers until the next write. Search never used them on the indexing node anyway.

Why do idle indexing shards use memory in Elasticsearch Serverless?

Elasticsearch Serverless stores Lucene commits and translogs in the object store, along with cluster state. Indexing nodes and search nodes are separate tiers. Durability no longer depends on replica shards, and relocating an indexing shard doesn’t copy files from node to node. The Symposium on Cloud Computing (SoCC) paper covers that architecture.

After the files left the node, Lucene's in-memory view of them stayed:

  • The standard IndexEngine holds an IndexWriter (even with an empty RAM buffer) and a DirectoryReader with its SegmentReaders opened so the writer can see existing documents for updates and deletes, in addition to version checks.

  • It kept per-index mappings and field infos. It also kept shard commit state.

Heap dumps of idle primary shards showed most of the retained memory in the engine's segment readers, on the order of a few megabytes each. The writer and mappings were a smaller slice. A hollow engine cut that footprint, and the node's retained heap fell with it.

On projects with heavy data stream usage and many backing indices, that’s the difference between the indexing node you need while idle and the one a live engine formula predicts.

Ingest autoscaling still sizes the indexing tier from memory formulas that assume a live engine. Hollow shards exist so we can stop assuming that, once those formulas catch up.

How hollow shards differ from fully materialized indexing shards

A fully materialized indexing shard runs an IndexEngine. A hollow shard runs a HollowIndexEngine over the same Lucene files, with no search or write path. Table 1 compares the two:

IndexEngine (fully materialized)

HollowIndexEngine (hollow)

Lucene IndexWriter

Open

None

Segment readers

Open

None

Accepts writes

Yes

No; ingestion is blocked until the shard is unhollowed

Local search

Theoretically possible

Not possible

Merges

Run on the indexing node

Require unhollowing first

Engine stats

Computed live

Served from stats stored in the hollow commit

Recovery

Translog replay and cache prewarm

No replay, no prewarm, no blob store LIST

Translog node ID in commit

Set

Empty (this is the hollow flag)

Table 1. IndexEngine versus HollowIndexEngine on an indexing node.

When we hollow, we flush everything to Lucene, commit, and upload to object storage with a hollow flag. We reuse an existing metadata field, the "translog node ID," and leave it empty. Empty means that this is a hollow commit and that recovery has no translog to replay, so it opens a HollowIndexEngine instead of the standard IndexEngine.

Having no readers makes the HollowIndexEngine a useful trip wire if some call site tries to search on the indexing tier. It also throws if indexing is attempted against it, but ingestion is gated before it reaches the engine, so we can unhollow the shard and reload an IndexEngine.

Search routing is unchanged; and queries continue to run on the search tier. Indexing routing still points at this primary. Since ingestion is blocked, a write queues until we unhollow.

When Elasticsearch Serverless hollows a shard

A shard can be hollowed only when all of these hold:

  • The index isn’t a system index.

  • It has been idle long enough for its type. (A regular index or the current write index of a data stream needs to be idle for one day by default, while older data stream backing indices need to be idle for 15 minutes.)

  • There’s no in-flight ingest or translog upload.

  • No merge is running or scheduled.

Why the first write to a hollow shard is slower

The first write is slower than a write against a warm engine because we unhollow first. We get less idle heap and faster shutdowns, and we accept a hitch on the first bulk to a cold backing index. On data streams that hitch almost never happens, because rolled-over backing indices aren’t the write index. They sit hollow until someone runs an update-by-query or a similar correction.

Hollowing a shard during  relocation

We only hollow a shard when it relocates. Relocations already block ingestion and already flush, so we reuse that window.

Figure 1. Hollowing an indexing shard during relocation. The source node A flushes a hollow commit to the object store. The target node B loads a hollow engine.

On the source node:

  • Ingestion is blocked (for the relocation).

  • If the shard can be hollowed, we force-flush a hollow commit and upload it. If it was already hollow, we skip the flush.

  • The target recovers that commit and sees the empty translog node ID (the hollow flag). It then loads a HollowIndexEngine.

  • We keep ingestion blocked on the target so the first write cannot sneak into the hollow engine.

Hollow recovery skips translog replay and cache prewarm because an idle shard doesn’t need them. If an indexing node restarts, a hollow commit on the object store comes back as a hollow engine.

How the first write unhollows a shard

The hollow engine cannot index. The first ingest, or a force-merge, has to put an IndexEngine back.

Figure 2. Unhollowing on the first ingest or force-merge. We flush a commit with a real translog node ID, recover it, prewarm, swap in an IndexEngine, and then release the ingestion block.

Incoming writes wait behind the ingestion block and trigger one asynchronous unhollowing that:

  • Runs the recovery work we skipped while hollow, including blob cache prewarming.

  • Closes the HollowIndexEngine and opens an IndexEngine, which is where Lucene's IndexWriter and DirectoryReader get created.

  • Skips translog replay. A hollow commit is clean; there’s nothing to replay.

  • Force-flushes a non-hollow commit, stamped with this node's translog ID, so a crash before the next regular commit upload still knows where the translog lives.

  • Releases the ingestion block. The waiting ingest continues against a real writer.

Reads don’t unhollow, and in general, they don’t come to the indexing tier. Real-time gets can be answered from the search tier. The data is already in the last hollow commit.

Force-merge is the other explicit materialization point. Serverless does not expose force-merge as a user-facing API, but we still use it internally; for example, on data stream backing indices. A merge needs an IndexWriter, so we unhollow and do the work. We let the shard go hollow again on a later relocation.

Swapping the engine on a live shard

The hardest part was an invariant that Elasticsearch has had for a long time; that is, once an IndexShard is started and live, the engine it opens is the engine it keeps.

There was no concept of changing the engine type under a started shard. Hollowing and unhollowing have to swap the engine on a live shard. Call sites that do getEngine().foo() assume the reference that they just got will still be open a few lines later. Two failure modes showed up immediately:

  1. A single engine call in the middle of a reset (for example, AlreadyClosedException on a real-time get).

  2. A call site that captures the engine and invokes it more than once (flush, block ingestion, and then flush again) across a swap.

We shipped a new resetEngine() on IndexShard that takes the shard and engine mutexes, plus a withEngine read/write lock so reset is the only writer. Getting there meant chasing deadlocks through refresh listeners, which now receive the reader they need without re-entering the engine, and putting withEngine guards on the call sites that assumed the engine couldn’t vanish. The result is fenced and a bit ugly.

Keeping stats and recovery off the Lucene readers

Telemetry would have reopened the readers that we had dropped. Doc stats wanted to open a searcher. After we hollow a shard, we compute the basic stats and store them in the hollow commit. We serve them from there, with no reading on the indexing node. 

Recovering a hollow commit used to LIST the blob store and prewarm cache chunks for files that the hollow engine will never read. We skip LIST now and keep the minimum file chunks inside the hollow commit. We skip prewarming unless we’re unhollowing.

Current limitations of hollow shards

There are two operational risks that we designed for and have not fully taken on yet: 

  • Autoscaling still counts hollow shards as full when it sizes the indexing tier, which is conservative. Later, we’ll charge them for the small footprint they actually have and keep a reservation for shards that might unhollow.

  • The related problem is a burst. After the autoscaling formulas tell the truth, a large update that wakes many backing indices at once could try to materialize more heap than the node has. We plan to throttle concurrent unhollowings against reserved heap.

Shutdown times and memory usage in production

In the first month after the rollout, fleet-wide indexing-node shutdown times improved at high percentiles by up to 30%, as Figure 3 shows. Unhollowings stay rare, while hundreds of thousands of shards have aged into a hollow state.

Figure 3. Fleet-wide indexing tier shutdown times before and after the rollout.

Indexing tier heap usage also dropped across the fleet, though we’re still working out how to publish this footprint to the autoscaler. The stability effect that we actually felt was shutdown time. Nodes finish leaving the fleet faster, which is what autoscaling needed.

On one project, shutdown times and indexing tier heap improved by up to 3x, as Figure 4 shows:

Figure 4. A project with 3x better shutdown times and heap usage of the indexing tier.

We also took heap dumps from that project. Before rollout, segment readers dominated the retained heap. After rollout, those types dropped off the top of the dump and the indexing tier scaled down to smaller pods.

Figure 5. Heap dump of a project after hollow shards, with Lucene readers gone from the top of retained heap.

What hollow shards mean for Elasticsearch Serverless projects

Thin indexing shards moved Lucene files off the local disk and onto the object store. Hollow shards move Lucene objects off the heap until a write needs them. A hollow commit is a batched compound commit (BCC) whose empty translog node ID means recover metadata, not a writer. It also uses a live engine swap, which IndexShard didn’t previously allow on a started shard, so the first bulk waits instead of racing the swap.

In production, high-percentile fleet shutdown times dropped by up to 30%, and we got back some of the indexing tier heap. The cost shows up on the rare first write to an idle shard or when we run a merge.

If you’re running Elasticsearch Serverless, some of your shards may already be hollow. Idle backing indices should relocate faster during upgrades and scale events, in line with the shutdown time drop in Figure 3, and they should sit cheaper on the indexing tier while they wait for the next write (or never get one).

Acknowledgments

Hollow shards were built by the Elasticsearch Distributed team. In addition to the authors, we would like to thank Francisco Fernández Castaño, Tanguy Leroux, Benjamin Lerer, Albert Zaharovits, Artem Prigoda, and Henning Andersen, who designed, implemented, and productionized the work. We would also like to thank the wider Elasticsearch engineering team.

Frequently Asked Questions

Do hollow shards affect search latency?

No. Search runs on the search tier from the object store commit and the blob cache. Hollowing only changes what the indexing node keeps in memory.

Will the first write to an old backing index be slower?

Yes. That write unhollows the shard (recover, prewarm, open IndexWriter). We’ve made that path as cheap as we can. Across a month of internal observability of unhollowings, their median was 218ms.

How helpful was this content?

Related Content

Elasticsearch Vector Database: Ship in minutes, scale affordably to hundreds of billions

Elasticsearch Vector Database: Ship in minutes, scale affordably to hundreds of billions

Dustin Coates
No more allocation delays: Decoupling snapshots from shard relocation in stateless Elasticsearch

No more allocation delays: Decoupling snapshots from shard relocation in stateless Elasticsearch

David Turner
Avoiding and Correcting Hotspots: How Elasticsearch Serverless Balances Shards

Avoiding and Correcting Hotspots: How Elasticsearch Serverless Balances Shards

Dianna Hohensee
Your AI agent doesn't need your API key: OAuth 2.1 for Elasticsearch MCP server authentication

Your AI agent doesn't need your API key: OAuth 2.1 for Elasticsearch MCP server authentication

Alex Chalkias

Ready to build state of the art search experiences?

Sufficiently advanced search isn’t achieved with the efforts of one. Elasticsearch is powered by data scientists, ML ops, engineers, and many more who are just as passionate about search as you are. Let’s connect and work together to build the magical search experience that will get you the results you want.