What is high cardinality?
What is high cardinality?
High cardinality means having so many unique label combinations in metrics that tracking each as a separate series creates scale problems for your time series database. Cardinality itself is the measure of uniqueness of values in a dataset, that is, how many distinct combinations exist. Labels—also called dimensions—are the metadata attached to a metric (for example, host, region, container_id.) Each unique combination is a separate series, and at scale those series are what drive memory exhaustion, runaway cost, and the data loss that comes from dropping labels just to stay under a limit.
Cardinality applies to any dataset. In a database table, a column of Boolean values has low cardinality: it can only hold true or false. A column of user IDs has high cardinality, with one distinct value for every user. In metrics, low cardinality means a label has a small, stable set of values, such as region or status_code, and high cardinality means the number of label combinations keeps growing. Cardinality is different from dimensionality, which is the number of labels a metric carries. Each label you add multiplies the possible combinations by the number of values that label can take. A cardinality explosion is what happens when that multiplication runs away: a sudden, sharp increase in the number of series, usually after a label with unbounded values, such as a user ID, is added to a metric.
What causes high cardinality
A metric is rarely just a number. For example, http_requests_total on its own is a single counter, but in practice it arrives with labels, which are key-value pairs attached to a metric that describe its context (host=web-01, region=us-east). In OpenTelemetry, the same key-value pairs are called attributes. Those labels are what make a metric useful, because they let you ask which host, which region, or which endpoint. They can also increase storage and processing costs.
The reason lies in how metrics are stored. A time series is a single named stream of timestamped values, and the database stores and indexes each one independently. The same metric name with three different sets of label values is therefore three separate series:
http_requests_total{host="web-01", region="us-east", endpoint="/checkout"}
http_requests_total{host="web-02", region="us-east", endpoint="/checkout"}
http_requests_total{host="web-02", region="eu-west", endpoint="/search"}
Each series holds its own sequence of timestamped values. Sampled every 15 seconds, the first one looks like this:
http_requests_total{host="web-01", region="us-east", endpoint="/checkout"}
10:00:00 1027
10:00:15 1034
10:00:30 1046
The number of possible series is the product of the label value counts. Each distinct set of label values creates a separate time series. So, if you have 1,000 hosts × 10 regions × 100 endpoints, that is 1 million series, all from one metric name. That figure is a ceiling, reached only if every host reports every endpoint in every region. Series exist only for label combinations that actually occur, so 1,000 hosts that each sit in one region produce 100,000 series.
The most common cause of runaway cardinality is using dynamic identifiers as metric labels: values that change with every session, request, or deployment, such as session IDs, request IDs, transaction hashes, unique user tokens, raw URL paths, temporary IP addresses, and pod names. Add one of them to a metric, and the total stops being bounded by the size of your infrastructure. It becomes bounded by your traffic.
The second factor is series churn: the constant creation and deletion of time series as Kubernetes containers come and go. A series does not stop costing you the moment its label values disappear. It still occupies the index, and a replacement series has already taken its place. A time series database that handles 1 million stable series efficiently may perform poorly when those series are replaced every few hours.
Benefits of high cardinality
High cardinality is often a sign that your telemetry is rich and detailed. The labels that multiply your series are the same ones that let you narrow a problem down to one host, one container, or one customer:
- Per-instance and per-user debugging: Detailed identifiers let you pinpoint the specific container that failed or follow one user's request across ten microservices.
- Anomaly detection: An unusual pattern in a single series stands out, where a broad aggregate would average it away.
- Faster root cause analysis: Every dimension you keep is a question you can ask during an incident without changing your instrumentation first.
The challenge is putting each value in the right data type. Bounded values such as region or service make good metric labels, while high-cardinality values such as user IDs belong in logs and traces.
Why high cardinality matters
High cardinality exhausts memory in systems that keep the index of active series in RAM, such as Prometheus, and raises the bill on platforms that charge per series.
Prometheus: the in-memory index. Prometheus keeps its full time series index in RAM—cardinality spikes cause out-of-memory (OOM) crashes. That in-memory index belongs to the head block, which holds roughly the last one to three hours of data. Older blocks store their indexes on disk and access them as needed. The head block holds an entry for every active series, next to an open chunk of recent samples for each one, so memory use is driven mainly by how many series are active and much less by how many samples you collect. That is why an instance can handle a high sample rate but run out of memory at a lower rate if the samples span too many series.
The design is fast, and it works well until cardinality spikes. When a deploy adds a new label, or an instrumentation change starts emitting request IDs as a dimension, the active series count climbs, and the head block grows with it. The result is an out-of-memory (OOM) crash—when Prometheus runs out of RAM because too many unique series are active simultaneously. These crashes tend to cluster around deploys, crash-looping pods, and autoscaling bursts, which is often exactly when you need the metrics most.
Per-series billing: the custom metrics bill. On platforms that bill per series, high cardinality can increase costs. Some pricing models charge per unique metric/label combination, so high-cardinality environments can generate millions of billable custom metrics. Every new combination of label values counts as another custom metric and adds to your billable usage. Per-host charges grow with your fleet, but label combinations can multiply far faster than hosts, so the custom metrics part of the bill can outpace the growth of your fleet.
How to manage high cardinality
Teams face a difficult tradeoff: reduce the detail in their metrics, risk running out of memory, or accept higher costs. These practices help keep cardinality under control, starting with the ones that remove the least detail:
- Avoid dynamic identifiers in metric labels entirely: Use bounded values instead, such as the route template /users/{id} in place of the raw path /users/12345 .
- Route high-variance context (user IDs, request IDs) to logs and traces instead: A request ID never belonged in a metric label. It is more useful on a log line or in a trace, where every request already carries its own ID.
- Set governance standards for instrumentation: Without shared standards, teams add labels individually until the stack breaks. Agree on which labels each metric may carry, and review new labels before they ship.
- Audit third-party integrations and collector defaults: They often emit high-cardinality labels you didn't configure. Check what exporters, agents, and OpenTelemetry collectors attach by default before the data reaches storage.
- Drop unused labels at ingest: Relabeling rules drop a label, or a whole series, before it is stored. The removed labels or series are no longer available for analysis.
- Use recording rules for dashboards: They precompute aggregates so dashboards can query a small set of summary series. They don't reduce the number of raw series ingested by the server that evaluates them. The saving comes when only the aggregates are forwarded to a global or long-term store, and there the per-host or per-endpoint detail is gone.
- Set per-target sample limits as a safety net: They protect the server by treating the entire scrape as failed when a target exceeds the limit. That trades a crash for a gap in your data.
- Use tag allowlists on platforms that bill per series: They restrict which tags stay queryable, which reduces costs but also limits the detail available for queries.
Preventing high-cardinality labels at the source costs you no useful information. Controls applied after the fact remove potentially useful infrastructure monitoring data before teams know which dimensions they'll need to investigate an incident.
High cardinality in practice: Kubernetes
High cardinality can come from any label whose values keep changing: user IDs, session IDs, request IDs, trace IDs, raw URL paths, IP addresses, container IDs, and pod names. It is especially common in Kubernetes environments. Containers in Kubernetes are ephemeral by design. Kubernetes creates and destroys containers continuously—each with a unique container ID that becomes a new label value. Pods managed by a Deployment get new names on every rollout, and every replacement Pod gets a new UID. Nodes are added and drained by a node autoscaler. Deployments, replica sets, and jobs all leave their own identifiers behind in the label set. None of this is misconfiguration. It is how the platform is designed to work.
Then there are per-pod metrics (CPU, memory, and network), each tracked across every unique pod/namespace/node combination. Several sources report pod-level series at once. The kubelet and its built-in cAdvisor expose resource usage for every container, while the kube-state-metrics add-on describes the state of every pod, deployment, and job as a series of its own.
The arithmetic works as it did above, with one difference: churn. Even at a conservative 50 series per pod, 5,000 pods is 250,000 active series at any given moment. The real figure is usually higher, because many container metrics repeat for every container, network interface, and disk device, but a count in that range is still manageable. The problem is that pod identifiers do not persist. A cluster that recycles most of its pods daily produces close to a fresh 250,000 the next day, and the day after that, so the number of distinct series accumulating across a retention window runs into the millions even while the count at any single instant still looks reasonable. This is how a medium-sized Kubernetes cluster can generate millions of unique time series without anyone doing anything wrong.
More frequent turnover makes the problem worse. A cluster running frequent deploys, CronJobs, or aggressive horizontal autoscaling replaces a large share of its pod identifiers every day. The faster that turnover, the more dead series are still sitting in the index when new ones arrive, and the higher the total climbs before old series age out. This is why high cardinality is a central challenge of modern cloud-native metrics and why Kubernetes monitoring is usually where teams first hit the ceiling.
How Elastic handles high cardinality
Elastic stores time series data in a columnar format on disk: each field is stored in its own column rather than row by row. Each backing index covers a set time range, and within it the values in every column are sorted by time series and then by timestamp. That lets the storage engine compress each column aggressively and read only the fields a query needs.
On disk does not mean slow. Elasticsearch relies on the operating system's filesystem cache, and Elastic's guidance is to leave at least half of a node's memory to that cache so that hot regions of the index stay in physical memory. The data you query most is still served from RAM.
There is no in-memory cardinality ceiling. Elastic's index lives on disk rather than in RAM, and Elasticsearch does not keep per-series state in memory that grows with cardinality, so adding new series does not drive up memory pressure on the nodes ingesting them.
Elasticsearch uses circuit breakers to stop a query whose estimated memory use would exceed a set share of the heap and returns an error instead. On Elastic, high cardinality costs you storage and query execution.
Time Series Data Streams (TSDS) is the Elastic feature that enables columnar TSDB storage—the mechanism behind the cardinality and efficiency advantages. Fields marked as dimensions form a generated series identifier; documents are routed by dimension so that every point in a series lands on the same shard; and segments are sorted by series and timestamp so that similar values sit next to each other and compress well. Combined with synthetic _source, the effect is substantial: in Elastic's benchmarks, metrics data stored in a TSDS used up to 70% less disk space than a regular data stream (about 30% from TSDS and about 40% from synthetic _source). Elastic's columnar metrics engine now stores OpenTelemetry metrics at 3.75 bytes per data point. That is up to 6.6x more efficient than earlier versions of TSDS, which needed 25 bytes, and up to 2.5x more efficient than Prometheus.
Downsampling reduces storage use for older data. A TSDS can downsample older data into fixed intervals. With the default method, it keeps the minimum, maximum, sum, and count of each gauge's values instead of every raw sample. It writes one document per series per interval and copies every dimension across, so each series survives with its labels intact. The series that mattered during last week's incident are still there after the raw samples are gone, at a coarser resolution and a fraction of the storage.
Moving to Elastic doesn't mean rebuilding your metrics pipeline. Elasticsearch accepts Prometheus Remote Write, so your existing scrape configs only need a new remote write target, and it ingests OpenTelemetry metrics natively over OTLP. PromQL works natively in Kibana, and a Prometheus-compatible API lets existing dashboards and other Prometheus clients query Elasticsearch directly, so your PromQL queries run unchanged.
Elastic's pricing is based on data volume or provisioned resources. The Elastic metrics platform does not price metrics per unique series or per custom metric. On Elastic Observability Serverless, metrics are billed per GB ingested and per GB retained; on Elastic Cloud Hosted, you pay for the resources you provision. Either way, cardinality growth shows up in the data volume you send or the capacity you provision, rather than as a line item that multiplies with every new label value. You can keep the labels you need (container_id, pod, customer_id, and the rest) and query on them, rather than deciding at ingest time which dimensions you can afford to keep.