<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Eugene Cistiakovas - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Eugene Cistiakovas - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/search-labs/author/eugene-cistiakovas</link>
    </image>
    <link>https://www.elastic.co/search-labs/author/eugene-cistiakovas</link>
    <atom:link href="https://www.elastic.co/search-labs/rss/author/eugene-cistiakovas.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Tue, 29 Sep 2026 21:43:37 GMT</lastBuildDate>
  <item>
    <title><![CDATA[How Elasticsearch Serverless hollow shards cut indexing-node shutdowns by 30%]]></title>
    <description><![CDATA[Idle indexing shards in Elasticsearch Serverless now drop their Lucene writers and segment readers from memory until the next write, freeing indexing-tier heap at the cost of a median 218ms wait on that first write.]]></description>
    <content:encoded><![CDATA[<p><a href="https://www.elastic.co/cloud/serverless">Elasticsearch Serverless</a> now unloads idle indexing shards from memory. We call them hollow shards; the shard stays allocated, but its<a href="https://lucene.apache.org/"> Lucene</a> <code>IndexWriter</code> and segment readers are gone until the next write arrives. Hollow shards build on<a href="https://www.elastic.co/search-labs/blog/thin-indexing-shards-elasticsearch-serverless"> thin indexing shards</a>, which already moved Lucene files off the local disk. In the first month in production, indexing-node shutdown times dropped by up to 30% at high percentiles across the fleet, and hundreds of thousands of shards went hollow. On one project with many idle data stream backing indices, shutdown time and indexing-tier heap improved by up to 3x. Search latency is unchanged, and the first write to a hollow shard waits a median 218ms while the shard reloads.</p><p>Thin shards write Lucene (the search library that Elasticsearch uses for indexing and search) segments locally and upload them to the object store as batched compound commits. Then they delete them from disk.</p><p>A thin shard that isn't ingesting and isn't expected to do so in the near future still holds an <code>IndexWriter</code>, segment readers, mappings, and the rest of its engine. At <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams">data stream</a> scale, those idle engines make up most of the indexing heap and add time to the relocations that we need for balancing and for shutting nodes down during autoscaling.</p><p>A <em>hollow indexing shard</em> is a thin allocated primary that keeps only the last commit's metadata and doesn't create its <code>IndexWriter</code> or segment readers until the next write. Search never used them on the indexing node anyway.</p><h2>Why do idle indexing shards use memory in Elasticsearch Serverless?</h2><p>Elasticsearch Serverless stores Lucene commits and translogs in the object store, along with cluster state. Indexing nodes and search nodes are separate tiers. Durability no longer depends on replica shards, and relocating an indexing shard doesn’t copy files from node to node. The <a href="https://www.elastic.co/search-labs/blog/elasticsearch-serverless-stateless-architecture">Symposium on Cloud Computing (SoCC) paper</a> covers that architecture.</p><p>After the files left the node, Lucene's in-memory view of them stayed:</p><ul><li><p>The standard <code>IndexEngine</code> holds an <code>IndexWriter</code> (even with an empty RAM buffer) and a <code>DirectoryReader</code> with its <code>SegmentReader</code>s opened so the writer can see existing documents for updates and deletes, in addition to version checks.</p></li><li><p>It kept per-index mappings and field infos. It also kept shard commit state.</p></li></ul><p>Heap dumps of idle primary shards showed most of the retained memory in the engine's segment readers, on the order of a few megabytes each. The writer and mappings were a smaller slice. A hollow engine cut that footprint, and the node's retained heap fell with it.</p><p>On projects with heavy data stream usage and many backing indices, that’s the difference between the indexing node you need while idle and the one a live engine formula predicts. </p><p><a href="https://www.elastic.co/search-labs/blog/elasticsearch-ingest-autoscaling">Ingest autoscaling</a> still sizes the indexing tier from memory formulas that assume a live engine. Hollow shards exist so we can stop assuming that, once those formulas catch up.</p><h2>How hollow shards differ from fully materialized indexing shards</h2><p>A fully materialized indexing shard runs an <code>IndexEngine</code>. A hollow shard runs a <code>HollowIndexEngine</code> over the same Lucene files, with no search or write path. Table 1 compares the two:</p><p></p><p>
</p><p><strong>IndexEngine (fully materialized)</strong></p><p><strong>HollowIndexEngine (hollow)</strong></p><p>Lucene <code>IndexWriter</code></p><p>Open</p><p>None</p><p>Segment readers</p><p>Open</p><p>None</p><p>Accepts writes</p><p>Yes</p><p>No; ingestion is blocked until the shard is unhollowed</p><p>Local search</p><p>Theoretically possible</p><p>Not possible</p><p>Merges</p><p>Run on the indexing node</p><p>Require unhollowing first</p><p>Engine stats</p><p>Computed live</p><p>Served from stats stored in the hollow commit</p><p>Recovery</p><p>Translog replay and cache prewarm</p><p>No replay, no prewarm, no blob store LIST</p><p>Translog node ID in commit</p><p>Set</p><p>Empty (this is the hollow flag)</p><p><em>Table 1. IndexEngine versus HollowIndexEngine on an indexing node.</em></p><p></p><p>When we hollow, we flush everything to Lucene, commit, and upload to object storage with a hollow flag. We reuse an existing metadata field, the "translog node ID," and leave it empty. <em>Empty</em> means that this is a hollow commit and that recovery has no translog to replay, so it opens a <code>HollowIndexEngine</code> instead of the standard <code>IndexEngine</code>.</p><p>Having no readers makes the <code>HollowIndexEngine</code> a useful trip wire if some call site tries to search on the indexing tier. It also throws if indexing is attempted against it, but ingestion is gated before it reaches the engine, so we can unhollow the shard and reload an <code>IndexEngine</code>.</p><p>Search routing is unchanged; and queries continue to run on the search tier. Indexing routing still points at this primary. Since ingestion is blocked, a write queues until we unhollow.</p><h3>When Elasticsearch Serverless hollows a shard</h3><p>A shard can be hollowed only when all of these hold:</p><ul><li><p>The index isn’t a system index.</p></li><li><p>It has been idle long enough for its type. (A regular index or the current write index of a data stream needs to be idle for one day by default, while older data stream backing indices need to be idle for 15 minutes.)</p></li><li><p>There’s no in-flight ingest or translog upload.</p></li><li><p>No merge is running or scheduled.</p></li></ul><h3>Why the first write to a hollow shard is slower</h3><p>The first write is slower than a write against a warm engine because we unhollow first. We get less idle heap and faster shutdowns, and we accept a hitch on the first bulk to a cold backing index. On data streams that hitch almost never happens, because rolled-over backing indices aren’t the write index. They sit hollow until someone runs an <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-update-by-query">update-by-query</a> or a similar correction.</p><h2>Hollowing a shard during  relocation</h2><p>We only hollow a shard when it relocates. Relocations already block ingestion and already flush, so we reuse that window.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb9f13cdc4d1a9ce0/6abbd0bab2d147043277178e/image8.jpg" alt="Diagram of hollowing a shard during shard relocation: a hollow commit in the object store loads as a HollowIndexEngine" /><p>On the source node:</p><ul><li><p>Ingestion is blocked (for the relocation).</p></li><li><p>If the shard can be hollowed, we force-flush a hollow commit and upload it. If it was already hollow, we skip the flush.</p></li><li><p>The target recovers that commit and sees the empty translog node ID (the hollow flag). It then loads a <code>HollowIndexEngine</code>.</p></li><li><p>We keep ingestion blocked on the target so the first write cannot sneak into the hollow engine.</p></li></ul><p>Hollow recovery skips translog replay and cache prewarm because an idle shard doesn’t need them. If an indexing node restarts, a hollow commit on the object store comes back as a hollow engine.</p><h2>How the first write unhollows a shard</h2><p>The hollow engine cannot index. The first ingest, or a force-merge, has to put an <code>IndexEngine</code> back.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcda5cef461180d8d/6abbd0d5d89724529109ad55/image1.jpg" alt="Diagram of the first write unhollowing an Elasticsearch Serverless shard, from HollowIndexEngine back to an IndexEngine" /><p>Incoming writes wait behind the ingestion block and trigger one asynchronous unhollowing that:</p><ul><li><p>Runs the recovery work we skipped while hollow, including blob cache prewarming.</p></li><li><p>Closes the <code>HollowIndexEngine</code> and opens an <code>IndexEngine</code>, which is where Lucene's <code>IndexWriter</code> and <code>DirectoryReader</code> get created.</p></li><li><p>Skips translog replay. A hollow commit is clean; there’s nothing to replay.</p></li><li><p>Force-flushes a non-hollow commit, stamped with this node's translog ID, so a crash before the next regular commit upload still knows where the translog lives.</p></li><li><p>Releases the ingestion block. The waiting ingest continues against a real writer.</p></li></ul><p>Reads don’t unhollow, and in general, they don’t come to the indexing tier. Real-time gets can be answered from the search tier. The data is already in the last hollow commit.</p><p>Force-merge is the other explicit materialization point. Serverless does not expose force-merge as a user-facing API, but we still use it internally; for example, on data stream backing indices. A merge needs an <code>IndexWriter</code>, so we unhollow and do the work. We let the shard go hollow again on a later relocation.</p><h2>Swapping the engine on a live shard</h2><p>The hardest part was an invariant that Elasticsearch has had for a long time; that is, once an <code>IndexShard</code> is started and live, the engine it opens is the engine it keeps.</p><p>There was no concept of changing the engine type under a started shard. Hollowing and unhollowing have to swap the engine on a live shard. Call sites that do <code>getEngine().foo()</code> assume the reference that they just got will still be open a few lines later. Two failure modes showed up immediately:</p><ol><li><p>A single engine call in the middle of a reset (for example, <code>AlreadyClosedException</code> on a real-time get).</p></li><li><p>A call site that captures the engine and invokes it more than once (flush, block ingestion, and then flush again) across a swap.</p></li></ol><p>We shipped a new <code>resetEngine()</code> on <code>IndexShard</code> that takes the shard and engine mutexes, plus a <code>withEngine</code> read/write lock so reset is the only writer. Getting there meant chasing deadlocks through refresh listeners, which now receive the reader they need without re-entering the engine, and putting <code>withEngine</code> guards on the call sites that assumed the engine couldn’t vanish. The result is fenced and a bit ugly.</p><h3>Keeping stats and recovery off the Lucene readers</h3><p>Telemetry would have reopened the readers that we had dropped. Doc stats wanted to open a searcher. After we hollow a shard, we compute the basic stats and store them in the hollow commit. We serve them from there, with no reading on the indexing node. </p><p>Recovering a hollow commit used to LIST the blob store and prewarm cache chunks for files that the hollow engine will never read. We skip LIST now and keep the minimum file chunks inside the hollow commit. We skip prewarming unless we’re unhollowing.</p><h2>Current limitations of hollow shards</h2><p>There are two operational risks that we designed for and have not fully taken on yet: </p><ul><li><p>Autoscaling still counts hollow shards as full when it sizes the indexing tier, which is conservative. Later, we’ll charge them for the small footprint they actually have and keep a reservation for shards that might unhollow.</p></li><li><p>The related problem is a burst. After the autoscaling formulas tell the truth, a large update that wakes many backing indices at once could try to materialize more heap than the node has. We plan to throttle concurrent unhollowings against reserved heap.</p></li></ul><h2>Shutdown times and memory usage in production</h2><p>In the first month after the rollout, fleet-wide indexing-node shutdown times improved at high percentiles by up to 30%, as Figure 3 shows. Unhollowings stay rare, while hundreds of thousands of shards have aged into a hollow state.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc964c2002c344f0e/6abbd1182693de74607120cd/image7.png" alt="Charts of hollow shards rising past 600,000 and p85 to p95 indexing node shutdown times dropping after the rollout" /><p>Indexing tier heap usage also dropped across the fleet, though we’re still working out how to publish this footprint to the autoscaler. The stability effect that we actually felt was shutdown time. Nodes finish leaving the fleet faster, which is what autoscaling needed.</p><p>On one project, shutdown times and indexing tier heap improved by up to 3x, as Figure 4 shows:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6a08d8d2718eb1a7/6abbd131d1edc65d4857043a/image3.png" alt="Dashboard of one Elasticsearch Serverless project where shutdown time and indexing tier heap usage dropped about 3x" /><p>We also took heap dumps from that project. Before rollout, segment readers dominated the retained heap. After rollout, those types dropped off the top of the dump and the indexing tier scaled down to smaller pods.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9f473cec04b38e6d/6abbd1437eb65f5810f0e2b2/image6.png" alt="Heap dumps before and after hollow shards: Lucene SegmentReaders dominate an 8GB pod and are gone from a 4GB pod" /><h2>What hollow shards mean for Elasticsearch Serverless projects</h2><p>Thin indexing shards moved Lucene files off the local disk and onto the object store. Hollow shards move Lucene objects off the heap until a write needs them. A hollow commit is a <a href="https://www.elastic.co/search-labs/blog/elasticsearch-refresh-costs-serverless">batched compound commit</a> (BCC) whose empty translog node ID means <em>recover metadata, not a writer</em>. It also uses a live engine swap, which <code>IndexShard</code> didn’t previously allow on a started shard, so the first bulk waits instead of racing the swap.</p><p>In production, high-percentile fleet shutdown times dropped by up to 30%, and we got back some of the indexing tier heap. The cost shows up on the rare first write to an idle shard or when we run a merge.</p><p>If you’re running Elasticsearch Serverless, some of your shards may already be hollow. Idle backing indices should relocate faster during upgrades and scale events, in line with the shutdown time drop in Figure 3, and they should sit cheaper on the indexing tier while they wait for the next write (or never get one).</p><h2>Acknowledgments</h2><p>Hollow shards were built by the Elasticsearch Distributed team. In addition to the authors, we would like to thank Francisco Fernández Castaño, Tanguy Leroux, Benjamin Lerer, Albert Zaharovits, Artem Prigoda, and Henning Andersen, who designed, implemented, and productionized the work. We would also like to thank the wider Elasticsearch engineering team.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-serverless-hollow-shards</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-serverless-hollow-shards</guid>
    <category><![CDATA[Elastic Cloud Serverless]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Operations]]></category>
    <dc:creator><![CDATA[Iraklis Psaroudakis,Eugene Cistiakovas]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt774d7c59a05d8932/6abbd030e80c0b53acaf0305/image2.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>