Configure data tiers for self-managed and Elastic Cloud on Kubernetes deployments

Whether you operate Elasticsearch on your own infrastructure or on Kubernetes with Elastic Cloud on Kubernetes, data tiers are expressed through each node’s data role. You choose which tiers the cluster offers by assigning the corresponding data_* roles to nodes or to ECK node sets.

The configuration steps for assigning data tier roles depend on your deployment type.

  1. For each node, decide which data tier or tiers it should participate in (for example data_hot, data_warm, data_cold, data_frozen, or data_content).
  2. Set node.roles in that node’s elasticsearch.yml to include the corresponding data_* roles (and any other roles the node should have, such as ingest or master).
  3. Restart the node or apply your configuration rollout process so the new roles take effect.

For example, the highest-performance nodes in a cluster might be assigned to both the hot and content tiers:

node.roles: ["data_hot", "data_content"]
		
Note

We recommend you use dedicated nodes in the frozen tier.

In Elastic Cloud on Kubernetes, a node set is a group of Elasticsearch pods that share one configuration. In the Elasticsearch manifest, set node.roles in that node set's config field (spec.nodeSets[].config). Use the same settings you would put in elasticsearch.yml on a self-managed host, and include a data_* role for each tier those pods should join.

This example assigns the hot and content tiers and the ingest role:

spec:
  nodeSets:
  - name: hot-content
    count: 3
    config:
      node.roles: ["data_hot", "data_content", "ingest"]
		

Some settings are managed by Elastic Cloud on Kubernetes; avoid overriding those. For the full mapping between manifest structure and Elasticsearch configuration, see Node configuration.

Note

On Elastic Cloud on Kubernetes, node set and scaling changes try to relocate shards from nodes that are removed, subject to allocation rules, capacity, and disk watermarks on the destination nodes. For more information, refer to the Elastic Cloud on Kubernetes documentation.

Follow this section when you need to remove a warm, cold, or frozen tier from a self-managed or Elastic Cloud on Kubernetes deployment. The hot and content tiers are required and cannot be removed. If you remove nodes that are assigned a data_hot or data_content role, ensure that the corresponding role remains assigned to other nodes.

The steps differ depending on whether the tier contains regular indices or searchable snapshot indices, which are common for cold or frozen tiers when using index lifecycle management (ILM).

If you plan to remove multiple tiers, remove them one at a time in this order: frozen, cold, then warm.

Important

Removing a data tier reduces the cluster's capacity. This can cause cluster instability, inaccessibility, or data loss if the remaining nodes cannot absorb the data from the removed tier.

Before proceeding:

  • Confirm that the remaining tiers have enough disk space, CPU, and memory to absorb the data and workload from the tier you are removing.
  • Review the disk watermarks and confirm that the nodes receiving the relocated shards have enough free disk space to remain below the low disk watermark.
  1. Identify which nodes belong to the data tier you want to remove:

    GET /_nodes?filter_path=nodes.*.name,nodes.*.ip,nodes.*.roles
    		

    Note the names of the nodes with the corresponding data_* role.

    Tip

    For Elastic Cloud on Kubernetes, also identify every nodeSet in your Elasticsearch manifest that has the data_* role associated with the tier you want to remove.

  2. Check whether the nodes in the tier you are removing hold shards from regular indices, searchable snapshot indices, or both:

    • Warm tier: This tier typically contains regular indices unless you have manually mounted searchable snapshots on it.

    • Cold tier: This tier can contain regular indices or fully mounted searchable snapshots. Check for standard ILM-managed searchable snapshot indices:

      GET /_cat/indices/restored-*?expand_wildcards=all
      		

      For each returned index, check its current data tier preference to determine whether it is on the tier you are removing.

      Exclude any fully mounted indices associated with the hot tier from the removal inventory. The hot tier is required and is not removed by this procedure.

    • Frozen tier: This tier contains only partially mounted searchable snapshots. Check for standard lifecycle-managed indices:

      GET /_cat/indices/partial-*,dlm-frozen-*?expand_wildcards=all
      		
    Note

    Manually mounted searchable snapshots might not use the standard restored-* or partial-* prefixes. If you mounted snapshots manually, adapt the index names or patterns in these requests to match your configuration.

  3. Review the ILM policies and index templates that can send data to the tier you are removing, and plan the changes required so that they no longer use the tier. This prevents newly created indices and future lifecycle transitions from targeting a tier that is no longer available.

    Depending on your configuration, plan to:

    • Remove or update ILM phases and actions that move indices to the tier.
    • Remove or move any searchable_snapshot action that mounts indices on the tier.
    • If you use custom allocation filters in policies or templates, remove or update those that target the tier.

    Make sure that your plan covers every affected policy and template.

    To learn more about ILM or shard allocation filtering, refer to Create your index lifecycle policy, Managing the index lifecycle, and Shard allocation filters.

When you have identified the nodes, determined which indices are on the tier, and planned the policy and template changes, continue with the procedure that matches that data:

This section explains how to vacate nodes in a data tier that contains searchable snapshot indices. How you proceed depends on the mount type:

Note

If any DLM-managed data stream uses frozen_after, remove this setting from the affected data stream lifecycles and index templates before removing the frozen tier. This prevents backing indices, including restored indices, from being converted to partially mounted searchable snapshots again.

  1. Apply the changes to ILM policies and index templates that you planned in Before you remove a data tier so that they no longer create or route searchable snapshot indices to the tier you want to remove. These changes prevent new searchable snapshots from appearing while you process the existing ones.

  2. For each partially mounted searchable snapshot, and for each fully mounted searchable snapshot that you do not want to keep mounted, select one of the following options:

    • Preserve the data as a regular index: Follow Restore searchable snapshot data to a regular index. Complete the restore, validation, alias or data stream update, and mounted index cleanup for one index before proceeding to the next.

    • Delete the data: Record the source snapshot details before deleting the index:

      GET /<searchable-snapshot-index-name>/_settings?filter_path=**.index.store.snapshot.snapshot_name,**.index.store.snapshot.repository_name&expand_wildcards=all
      DELETE /<searchable-snapshot-index-name>
      		

      If you no longer need the source snapshot, delete it after confirming that it contains no other data you need and that no other mounted index in this or another cluster depends on it:

      Warning

      After you delete the mounted index, deleting its source snapshot permanently removes the data if no other copy exists. Keep the source snapshot if you might need to restore the data later.

      DELETE /_snapshot/<snapshot_repository_name>/<searchable_snapshot_name>
      		

After processing all searchable snapshots, continue based on what remains on the tier:

Use this section to update shard allocation rules for regular indices before you remove the tier. Follow the same steps for fully mounted searchable snapshots that you want to keep mounted. Those snapshots use the same shard allocation rules as regular indices.

  1. If you have not already done so, apply the changes to ILM policies and index templates that you planned in Before you remove a data tier. These changes prevent newly created indices and future lifecycle transitions from targeting the tier. They do not move indices already allocated there. The remaining steps update those indices and relocate their shards.

    Warning

    Temporarily stopping ILM can prevent lifecycle transitions while you update the cluster configuration, but it affects every ILM-managed index in the cluster. It pauses actions such as rollover, migration, and deletion. On clusters with sustained ingestion, a long pause can cause indices on the hot tier to grow until the tier runs out of disk space.

    Keep ILM running unless you understand the effect on your workload. If you stop it, monitor the hot tier and restart ILM as soon as possible. Stopping ILM does not replace updating policies, templates, and index allocation settings.

  2. Determine which shards are allocated to the nodes you want to remove.

    GET /_cat/shards?v&h=index,shard,prirep,state,node
    		

    Filter the output by the node names you identified in Before you remove a data tier.

  3. Check and update index allocation rules.

    ILM and manual index configurations can use different index-level shard allocation filters to control shard placement. For every index that has shards on the nodes you are removing, check its allocation settings and complete the applicable steps:

    GET /my-index/_settings
    		
    1. Update _tier_preference-based rules.

      Data tier-based ILM policies use index.routing.allocation.include._tier_preference to express shard placement as an ordered list of preferred tiers. Elasticsearch allocates shards to the first tier in the list that has nodes in the cluster and considers later tiers only when none of the preceding tiers have any nodes.

      Indices using this method have settings similar to the following example:

      {
      ...
          "routing": {
              "allocation": {
                  "include": {
                      "_tier_preference": "data_warm,data_hot"
                  }
              }
          }
      ...
      }
      		
      1. The example represents an index in the warm tier.

      Before manually vacating the nodes, update _tier_preference so that the tier where you want the data to move is the first available tier in the list. This change makes the destination tier preferred and starts relocating the shards before the nodes are removed.

      Update the setting based on where you want to move the data:

      • To move the data to an existing fallback tier, remove the tier being removed from the list. For example, when removing the warm tier, change data_warm,data_hot to data_hot.
      • To move the data to a later lifecycle tier, add that tier before the tier being removed. For example, when removing the warm tier, change data_warm,data_hot to data_cold,data_warm,data_hot.

      The following example moves data from warm to cold:

      PUT /my-index/_settings
      {
          "routing": {
            "allocation": {
              "include": {
                  "_tier_preference": "data_cold,data_warm,data_hot"
              }
            }
          }
      }
      		
      1. You can also use data_cold,data_hot. Both values move the data to cold, but omitting data_warm removes that tier from the fallback sequence.
      Note

      Do not use the frozen tier as a fallback for regular indices or fully mounted searchable snapshots. It is reserved for partially mounted searchable snapshots.

    2. Review custom allocation rules.

      Some custom configurations use index-level shard allocation filters in addition to or instead of _tier_preference. These filters use require, include, or exclude rules with built-in or custom node attributes to control shard placement.

      For example, the following settings use a custom data node attribute to require warm nodes:

      {
      ...
          "routing": {
              "allocation": {
                  "require": {
                      "data": "warm"
                  }
              }
          }
      ...
      }
      		

      A require rule is a hard constraint. If no nodes match it, the shard remains unassigned. To remove this requirement:

      PUT /my-index/_settings
      {
        "index.routing.allocation.require.data": null
      }
      		
      1. You can update the rule to target the destination nodes instead of removing it.

      For each affected index, update or remove the custom filters that prevent allocation to the destination tier.

      The following example removes all _name-based allocation filters from an index:

      PUT /my-index/_settings
      {
        "index.routing.allocation.require._name": null,
        "index.routing.allocation.include._name": null,
        "index.routing.allocation.exclude._name": null
      }
      		

      Removing a custom filter does not necessarily start relocation if the current nodes remain eligible. The manual vacate in the following step forces any remaining shards to move.

  4. Vacate the nodes manually.

    Note

    On Elastic Cloud on Kubernetes, removing a nodeSet from the Elasticsearch manifest can migrate data away from its nodes before removing the underlying StatefulSet, as described in Cluster upgrade patterns. This procedure uses a manual vacate so that you can verify the nodes are empty before removing the nodeSet.

    Exclude the nodes from shard allocation by name. Elasticsearch then relocates their remaining shards to other eligible nodes:

    PUT /_cluster/settings
    {
      "persistent": {
        "cluster.routing.allocation.exclude._name": "<node-name-1>,<node-name-2>"
      }
    }
    		
    1. If _name exclusions are already configured, include their existing values in the comma-separated list to preserve them.
    Important

    Wait until GET /_cat/allocation?v=true&s=node shows that no shards remain on those nodes before proceeding. Updating settings starts the relocation process, but you must wait until shard allocation and recovery finish. If shards stay on the original tier, use the cluster allocation explain API to determine the cause. Refer to Using the cluster allocation API for troubleshooting for common examples. Common causes include disk watermarks or index.routing.allocation.total_shards_per_node limit reached on the destination nodes.

After the nodes are empty, continue to Remove the tier nodes.

After completing every applicable vacate procedure, follow these steps to remove the empty nodes from the tier.

  1. Confirm that no shards remain on the nodes you want to remove:

    GET /_cat/allocation?v=true&s=node
    		

    Do not continue until the nodes report no shards. If shards remain, complete the applicable vacate procedure and use the cluster allocation explain API to identify any allocation constraints.

  2. Remove the nodes.

    Stop the Elasticsearch service on each node to be removed and decommission the host. For step-by-step instructions, refer to Add or remove Elasticsearch nodes.

    Remove every nodeSet associated with the tier from your Elasticsearch manifest, or set each count to 0. If an ElasticsearchAutoscaler policy manages any of these nodeSets, remove the matching policy before applying this change. Otherwise, autoscaling might change the nodeSet counts while you complete this procedure. Refer to Autoscaling in ECK.

    Elastic Cloud on Kubernetes safely stops the pods after you have vacated their shards.

  3. Wait until GET /_cat/nodes?v shows no nodes from the removed tier remaining in the cluster.

  4. If you used the manual vacate, remove the deleted node names from the exclusion rule only after the nodes have left the cluster. Restore any _name exclusions that existed before the vacate. If none existed, clear the setting:

    PUT /_cluster/settings
    {
      "persistent": {
        "cluster.routing.allocation.exclude._name": null
      }
    }
    		
  5. Confirm that GET /_cluster/health reports green.

  6. Verify that ILM is running and that no indices report errors related to the removed tier:

    GET /_ilm/status
    GET /_all/_ilm/explain?human=true&expand_wildcards=all&only_errors=true
    		

    Confirm that operation_mode is RUNNING. Investigate any reported errors and verify that no policy still attempts to allocate data to the removed tier.

    For indices in the ERROR step, resolve the underlying cause first. You can then force ILM to retry the failed step immediately:

    POST /<affected-indexes>/_ilm/retry
    		

    For guidance, refer to Fix ILM errors.