What are metric types? Counter, gauge, histogram, and summary

What are metric types? Counter, gauge, histogram, and summary

Every metric your application exposes has a type, and that type determines which questions you can answer later. A counter lets you chart request rates and alert when errors spike. A summary gives you accurate latency percentiles for one instance but can't tell you the p99 of a whole service. There are four standard types: counter, gauge, histogram, and summary. They are the core metric types in Prometheus, and OpenTelemetry uses a closely related set.

A metric is a number that your application keeps in memory while it runs, like the requests it has served or the connections it has open. You define each metric in the application's instrumentation, using a Prometheus client library or an OpenTelemetry SDK, and give it a name, such as http_requests_total, and a type. A monitoring system records the current value at a regular interval (for example, every 15 seconds) and stores it with a timestamp, so it sees only those samples.


Why metric types matter

You choose the type once, in the instrumentation, and every query you write later depends on it. Count errors with a counter, and the difference between any two samples includes every error in between, even a burst that lasts two seconds. Report the same errors as a gauge that holds the current errors per second, and a burst that starts and ends between two samples never reaches your charts. Changing the type later means changing the instrumentation, redeploying, and rewriting the queries behind your dashboards and alerts.

Metric types are separate from chart types, even where the names overlap: a gauge metric is often plotted as a line chart, and a histogram metric often appears on a dashboard as a latency percentile over time.


Counter: tracking totals that only increase

A counter is a metric that only increases—tracking cumulative totals like http_requests_total, errors_total, or bytes_sent. It never decreases unless the process restarts. The application adds 1 for each request it serves or the size of each response it sends.

At each collection, the monitoring system stores the counter's current total:

10:00:00   48210
10:00:05   48230
10:00:10   48250

The total counts requests since the process started. To measure the current request rate, compare samples over a time interval. The difference between samples gives you that: here, 40 requests in 10 seconds, or 4 requests per second. To get a rate, use rate() in PromQL, which makes this calculation over a time window, such as the last five minutes:

rate(http_requests_total[5m])

The result is a requests-per-second series that you can chart or alert on. PromQL also has irate() , which calculates the rate from only the last two samples in the window. Use it to graph fast-moving counters where you want to see short spikes, and use rate() for alerts.

  • Key property: After a restart, the counter starts again from zero. rate() reads a drop in the value as a restart, so it doesn't return a negative result.
  • When to use it: Any cumulative count, such as requests, errors, bytes, completed tasks, or cache misses. If the value can go down, use a gauge.

Gauge: tracking current state

A gauge measures a value that can increase or decrease—current CPU percentage, memory in use, active database connections. It is a snapshot at the moment of collection. Unlike a counter, which keeps a running total, a gauge reports the current level, such as the number of connections open right now. The application sets it to the current value or moves it up and down as connections open and close. Typical gauges are cpu_usage_percent , memory_used_bytes , and active_connections .

At each collection, the monitoring system stores the gauge's current value:


10:00:00   212
10:00:15   230
10:00:30   198

Each sample is a complete measurement, so there is no rate to calculate, and you can chart a gauge as it is. Gauge values can be averaged, min/maxed, and compared directly: the average CPU usage over the last 10 minutes, the busiest instance of each service, or every instance that uses more than 90% of its memory limit.

  • Key property: Only the value at each collection is recorded, so a two-second CPU spike can fall between two samples taken 15 seconds apart.
  • When to use it: Current state measurements, such as resource usage, queue depth, pool size, or temperature. They tell you the current load on a service, how full a buffer or queue is, and how much capacity is left. If you are counting events, use a counter.

Histogram: measuring distributions and latency

A histogram records how values distribute across predefined buckets—for example, request latency grouped as under 50 ms, 50–200 ms, 200–500 ms, and over 500 ms. It stores a count for each bucket instead of a single value. Percentiles are calculated at query time.

An application that serves thousands of requests between two collections can't publish the duration of every request, so a histogram counts them instead. Each time a request finishes, the instrumentation in your application adds 1 to the bucket that its duration falls in, and it keeps a count of all requests and the sum of their durations. After 10,000 requests, the bucket counts are

under 50ms     8,200
50–200ms       1,200
200–500ms        500
over 500ms       100

What you usually want from a histogram is a percentile: the 95th percentile of latency, or p95, is the duration that 95% of requests stay under. Here, 9,400 requests finished in 200ms or less, so the 9,500th falls in the third bucket, between 200ms and 500ms. The query estimates its position inside that bucket and returns about 260ms. In PromQL, histogram_quantile() makes this calculation:

histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))

The query takes the rate of each bucket, adds the buckets up across all instances (le is the label that holds each bucket's upper bound), and estimates the percentile from the result. Prometheus publishes the counts cumulatively, so its 200ms bucket reads 9,400: every request that took 200ms or less. Because bucket counts from different instances can be added together, one histogram gives you percentiles for a whole service.

  • Key property: Buckets are defined at instrumentation time. A percentile is estimated inside one bucket, so its accuracy depends on the bucket boundaries.
  • Cardinality: Prometheus tracks each bucket of a histogram in its own series, so the latency histogram above is six series: four buckets, a count, and a sum. Every label you add multiplies those series by the number of values the label takes, so a histogram with many labels can quickly lead to high cardinality.
  • When to use it: Latency, request size, and anything else where the distribution matters more than individual values.

Summary: pre-calculated percentiles

A summary pre-calculates quantiles (p50, p95, p99) at the client before data is sent. They are accurate for a single instance but cannot be aggregated across multiple instances. A histogram leaves the same calculation to the query instead. A quantile is a percentile written as a fraction, so the 0.95 quantile is p95.

The client is the instrumentation library inside your application. It tracks the most recent observations, calculates the percentiles itself, and publishes the results alongside a count and a sum. For rpc_duration_seconds, that looks like this:

p50   0.012 seconds
p95   0.087 seconds
p99   0.21 seconds

No calculation is needed at query time. You chart the published p95 as it is, as you would a gauge.

Each instance calculates its percentiles from its own traffic, and percentiles can't be combined afterward: the average of ten p99 values is not the p99 of the service. For multi-instance percentiles, use a histogram instead.

  • Key property: The quantiles are chosen in the instrumentation and calculated by the client, so you can't ask for a different percentile at query time.
  • When to use it: Accurate percentiles for a single instance when you can't predict the range of values well enough to choose histogram bucket boundaries.

How to choose the right metric type

Most services use more than one type. A web service might count requests and errors with counters, report open connections with a gauge, and measure latency with a histogram. The table compares the four types:

TypeWhat it measuresAggregatable across instancesBest forExample
CounterA cumulative total that only increasesYes.Requests, errors, bytes, completed taskshttp_requests_total
GaugeA current value that goes up and downYes.Resource usage, queue depth, connectionsmemory_used_bytes
HistogramA distribution of values, counted in bucketsYes.Latency and size percentiles across a servicehttp_request_duration_seconds
SummaryQuantiles calculated by the clientNoAccurate percentiles for one instancerpc_duration_seconds

 

Two decisions cover most cases:

  • Counter vs. gauge: Pick a counter when the value only accumulates, and query its rate. Pick a gauge when the value can go down or you care about its current level.
  • Histogram vs. summary: Pick a histogram when you run more than one instance or may want a different percentile later. Pick a summary for accurate single-instance percentiles when you can't choose bucket boundaries in advance.

Metric types in OpenTelemetry

OpenTelemetry has its own type system because it's designed to be vendor neutral, so the same instrumentation can export to any backend. With an OpenTelemetry SDK, you record measurements through instruments, and the SDK exports them as one of four metric types, each of which maps to a Prometheus type:

  • Sum: A running total. A Sum that only increases maps to a Counter. A Sum that can also go down, which comes from an UpDownCounter instrument, maps to a Gauge.
  • Gauge: The current value at collection time, which maps to a Gauge.
  • Histogram: Counts in buckets with fixed boundaries, plus a sum and a count. It maps to a Histogram.
  • ExponentialHistogram: A histogram whose bucket boundaries follow an exponential scale and adjust to the data, which gives you higher resolution with no buckets to choose in advance. It maps to a native histogram, the newer Prometheus histogram format.

OpenTelemetry SDKs don't produce summaries. A Summary type exists so that summaries from Prometheus clients can pass through an OpenTelemetry pipeline.

Some monitoring systems also give each metric a kind alongside its value type, which is the data type of each point, such as an integer or a distribution. The kind says how points relate over time: a gauge is a current value, a cumulative metric is a running total since a fixed start time, and a delta metric holds only the change since the previous point. OpenTelemetry calls this choice temporality: Sums and both histogram types are either cumulative, like Prometheus counters and histograms, or delta.

When the destination is Elasticsearch, you don't build this mapping yourself:

EDOT (Elastic Distribution of OpenTelemetry) translates OTel metric types to Elasticsearch's storage format, preserving Sum, Gauge, Histogram, and ExponentialHistogram semantics.

A Sum that only increases is stored as a counter, a Gauge as a gauge, and both histogram types as histogram fields that return any percentile at query time.


How Elasticsearch stores all metric types

Elasticsearch is a search and analytics engine that also works as a time series database, so your metrics live in the same place as the logs and traces you search when something breaks. It handles all four metric types from either data model. Your instrumentation stays as it is.

  • OpenTelemetry: EDOT ingests OTel metric types natively. Your SDKs and Collectors send their data over OTLP, the OpenTelemetry protocol, and Elasticsearch stores exponential histograms in a native field type.
  • Prometheus: Prometheus Remote Write ingests Prometheus metric types. You add Elasticsearch as a remote_write destination, and counters, gauges, histogram buckets, and summary quantiles arrive as the same series that Prometheus scrapes, with their labels stored as dimensions.

Both paths write to time series data streams, the columnar storage that Elasticsearch uses for metrics. Each series is stored as a counter or a gauge: OpenTelemetry data carries its type, and Remote Write data is typed by the Prometheus naming conventions, such as the _total suffix on counters. Rate calculations then apply only to counters and account for resets.

For querying, ES|QL and PromQL are both supported. PromQL runs natively, so the rate() and histogram_quantile() queries above work as written. ES|QL adds time series aggregations and percentiles on histogram fields, and the same language queries your logs and traces.

Both languages run in Kibana, in Discover, in dashboard panels, and in alert rules. You can display a request rate as a line chart, show a gauge as a single number, or plot latency percentiles together on one chart. That gives you one place for metrics monitoring, whichever type and data model your services use.

For storage efficiency and query speed, see the Elasticsearch metrics benchmarks.