<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Felix Barnsteiner - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Felix Barnsteiner - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/search-labs/author/felix-barnsteiner</link>
    </image>
    <link>https://www.elastic.co/search-labs/author/felix-barnsteiner</link>
    <atom:link href="https://www.elastic.co/search-labs/rss/author/felix-barnsteiner.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Tue, 29 Sep 2026 15:27:56 GMT</lastBuildDate>
  <item>
    <title><![CDATA[How we built PromQL into Elasticsearch]]></title>
    <description><![CDATA[PromQL runs on the same Elasticsearch compute engine as ES|QL, with no plugin and no separate process to operate. Getting there meant changing how the engine evaluates time windows and builds grouping keys.]]></description>
    <content:encoded><![CDATA[<p>More than 80% of the Prometheus Query Language (PromQL) queries in our real-world corpus run on Elasticsearch without modification. Elasticsearch 9.5 makes the PromQL and the Prometheus-compatible API generally available (GA), so you can ingest Prometheus metrics with<a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/tsds-ingest-prometheus-remote-write"> remote write</a> and query them through the<a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api"> Prometheus HTTP APIs</a> or the<a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/promql"> PROMQL</a> command in Elasticsearch Query Language (ES|QL).</p><p>PromQL compiles to the same compute engine that runs ES|QL and inherits its planner and distributed execution, along with its release process. We didn’t build a second engine for this, and there’s no plugin to install.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9e310070ec244c5d/6aa3928e224d356c5e0dd18e/1.png" alt="PromQL compatibility in Elasticsearch rising from zero to 80% between 9.4 Tech Preview and 9.5 GA" /><p>This post is about how we built it.</p><p>Key takeaways:</p><ul><li><p><strong>One engine:</strong> The implementation combines Elasticsearch’s mature distributed planning, storage, and testing infrastructure with its newer compute engine, which provides a columnar execution runtime. This lets PromQL reuse proven Elasticsearch capabilities while executing through a modern, native vectorized pipeline rather than introducing a separate runtime.</p></li><li><p><strong>One server:</strong> Elasticsearch implements the Prometheus remote write and query APIs directly, so Prometheus-compatible ingest and queries run without any additional plugins.</p></li><li><p><strong>Engineered for efficiency:</strong> Supporting PromQL required new engine primitives for range-aligned evaluation grids, backward-looking windows, dynamic label grouping, pipeline result reshaping, and compact wide aggregation keys. These primitives allow PromQL queries to execute efficiently end to end, with the relevant semantics implemented directly in the compute engine rather than through external post-processing.</p></li><li><p><strong>Compatibility measured in real use:</strong> In addition to Prometheus compliance tests, we built a differential-testing and quality-control pipeline over 2,000 PromQL queries collected from public repositories. </p></li></ul><p>Read more:</p><ul><li><p><a href="https://www.elastic.co/observability-labs/blog/elasticsearch-supports-promql">Query Prometheus Metrics in Elasticsearch with PromQL</a></p></li><li><p><a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch">Ship Prometheus Metrics to Elasticsearch with Remote Write</a></p></li><li><p><a href="https://www.elastic.co/search-labs/blog/elasticsearch-native-prometheus-api">Bringing Fire to Elasticsearch: Adding Native Prometheus APIs</a></p></li></ul><h2><strong>Why run PromQL on Elasticsearch</strong></h2><p>Many teams already store logs and traces in Elasticsearch while running metrics in Prometheus or another dedicated metrics back end.</p><p>Prometheus and its ecosystem are strong and widely adopted, but large deployments can also bring operational sprawl and scaling challenges, along with limited retention. </p><p>So we set out to combine Elastic’s highly optimized time-series database (TSDB) with a best-in-class metrics ecosystem. The result is a smaller observability stack, with fewer systems to operate and metrics storage that scales horizontally and supports long-term retention.</p><h2><strong>One engine: PromQL and ES|QL share the same compute engine</strong></h2><p>We made an early architectural decision not to run a separate PromQL engine next to Elasticsearch.</p><p>PromQL is instead another front end to the Elasticsearch compute engine.</p><p>This puts PromQL in the normal Elasticsearch development lifecycle. It uses the same planner, distributed execution engine, testing infrastructure, and release process as Elasticsearch itself.</p><p>To learn more about Elasticsearch’s query engine, check <a href="https://www.elastic.co/search-labs/blog/elasticsearch-columnar-metrics-engine-30x-faster-prometheus">our blog</a>.</p><p>Like ES|QL’s time series queries that use the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts"><code>TS</code></a> source command, PromQL is translated into a highly optimized query plan and executed across the cluster. Nodes process columnar batches through vectorized operators, while partial results move through exchanges until the final result is assembled. </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcf3a413e89209805/6aa392b38406d9813bcaa1d8/2.png" alt="PromQL and ES|QL frontends feed one shared Elasticsearch planner and columnar execution DAG across shards" /><p>This also means that PromQL and ES|QL operate over the same execution engine and time-series data. ES|QL can additionally extend a PromQL computation with post-processing that PromQL doesn’t support, such as lookup joins and inline aggregations.</p><p>For example, assume Prometheus request counters are stored in <code>metrics-*</code> and keyed by the <code>service</code> label. A lookup index, <code>service_registry</code>, maps each instance to its owning team and environment and to its service tier:</p><p></p><p></p><p>This architecture requires the execution engine to support PromQL semantics natively and efficiently rather than ES|QL syntax sugar. The following sections describe the changes and new execution primitives we introduced to achieve that.</p><h2><strong>One server: Prometheus remote write and HTTP API built into Elasticsearch</strong></h2><p>Query execution is only half of the story. The Prometheus ecosystem also expects familiar ingest and query APIs.</p><p>Prometheus protocols are the de facto standard for everything metrics in almost every team’s observability stack. So we built the HTTP API directly in the Elasticsearch server, which eliminated the need for a third component and tightened integration stability and performance.</p><p>On the ingest side, we <a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">added</a> an endpoint for the <a href="https://prometheus.io/docs/specs/prw/remote_write_spec/">Prometheus remote write</a> protocol. It accepts Snappy-compressed Protocol Buffer messages, maps labels to TSDS dimensions, maps the metric name/value into metric fields, infers counter versus gauge mappings, and writes directly into TSDS. The built-in template is dynamic, so users don’t have to predeclare every Prometheus label or metric.</p><p>On the query side, Elasticsearch <a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">exposes</a> Prometheus query APIs. A request enters through the Prometheus endpoint and is executed in the compute engine.</p><h2><strong>Running PromQL efficiently in a columnar engine</strong></h2><p>Sharing an execution engine doesn’t mean treating PromQL as syntax sugar over ES|QL. PromQL has different time, grouping, and response semantics, as well as workload characteristics that matter at scale. Supporting it efficiently required extending the compute engine rather than compensating in the API layer.</p><h3><strong>PromQL time grids: Aligning evaluation steps with TSTEP</strong></h3><p>Time-series query engines are optimized for grouping over time.</p><p>Elasticsearch normally groups timestamps with <code>TBUCKET(...)</code>, which truncates each timestamp to a fixed interval boundary. Truncation is cheap and produces deterministic bucket boundaries. It also makes intermediate results easier to reuse.</p><p>Prometheus defines evaluation points differently. For a range query, timestamps are laid out as fixed steps anchored to the query range, rather than derived by truncating each sample timestamp. Two queries with the same step but different range boundaries can therefore produce different evaluation grids.</p><p>To preserve these semantics, we introduced <code>TSTEP(...)</code>, which derives its grouping grid from the query range and step,  rather than truncating timestamps to globally aligned boundaries.</p><p>PromQL uses <code>TSTEP(...)</code> internally, preserving Prometheus timestamp semantics while still lowering the operation to a native Elasticsearch execution primitive.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt437155bca2029671/6aa3937927a5312436dcbf38/3.png" alt="TSTEP vs TBUCKET in Elasticsearch: PromQL step grid anchored to query start, TBUCKET to fixed boundaries" /><p>Some might see this as a simple problem, but the nuances matter. </p><p>Take, for example, a query that finds a 5m rolling average of a metric: </p><p></p><p></p><p>At evaluation time <code>T</code>, the result represents the average over the preceding five-minute range: </p><p><code>(T - 5m, T]</code></p><p>When the query is executed with a five-minute step, each output value is labeled with the upper end of its corresponding five-minute window.</p><p>Elasticsearch previously lacked this semantic and supported only forward-looking window aggregation functions:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt97aac2fa95cf657d/6aa393b71ade64390142ee2e/4.png" alt="Forward-looking window aggregation where each bucket covers the interval from timestamp T to T plus W" /><p>We rewrote the window-evaluation path so that both ES|QL and PromQL use a common backward-looking windowing implementation:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1a12ceb27d9ab95b/6aa393dcf08ee141fe856719/5.png" alt="Backward-looking PromQL window covering T minus W to T, the range used by rate and avg_over_time" /><p>Together, <code>TSTEP(...)</code> and backward-looking windows preserve the two time semantics that matter for PromQL range evaluation.</p><h3><strong>Dynamic label grouping: How PromQL </strong><strong><code>without()</code></strong><strong> resolves at runtime</strong></h3><p>Time grids determine when a PromQL expression is evaluated. Aggregation determines which input series are combined and which labels identify each output series.</p><p>For most analytical query engines, that identity is known when the query is planned. The planner can allocate grouping columns and choose an aggregation strategy. It also carries a fixed key through the execution pipeline.</p><p>That’s how ES|QL works:</p><p></p><p></p><p>The output series are grouped by an explicit key <code>(cluster, namespace)</code>.</p><p>PromQL can express the same operation in the opposite direction:</p><p></p><p></p><p>Now we know which dimensions <em>not</em> to use. We don’t necessarily know the full grouping key until the query is executed. This is a small language difference with significant execution consequences. </p><p>One possible implementation is to discover every label used by the metric, subtract <code>instance</code> and <code>pod</code>, and rewrite the expression into an ordinary <code>by(...)</code> aggregation. That adds a <a href="https://www.elastic.co/search-labs/blog/esql-metrics-info-ts-info-time-series-catalog">discovery phase</a> before planning. It also becomes inefficient for high-dimensional metrics where only a subset of all possible dimensions may have useful value in a particular series. Most queries need only a small subset of the available dimensions, so carrying the entire dimension universe as an aggregation key wastes memory and adds bookkeeping overhead.</p><p>We instead extended the time-series execution path with dynamic grouping columns. The engine loads dimensions as the series are read and applies the exclusions per time series. This avoids making the grouping schema a prerequisite for planning and avoids carrying a large sparse set of grouping columns through aggregation.</p><h3><strong>Dimension packing: Keeping wide PromQL grouping keys cheap</strong></h3><p>The <a href="https://prometheus.io/docs/practices/rules/#aggregation">idiomatic</a> way of writing PromQL aggregations involves heavy use of <code>without(...)</code> over <code>by(...)</code>:</p><p></p><p></p><p>Excluding labels rather than explicitly listing them makes dashboards and alerts resilient to schema evolution. If a new label is added to the metric, the query continues to preserve it unless it’s explicitly excluded.</p><p>The consequence for the engine is that effective grouping keys can be wide. Many of those labels often have low cardinality, yet each still participates in every aggregation stage.</p><p>In a columnar engine, each grouping column is normally represented as a separate vector. Ten grouping labels therefore mean 10 vectors flowing through every aggregation operator:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2f491d0d9fdb82f2/6aa3940027a53105ccdcbf3c/6.png" alt="Elasticsearch columnar page: rows split into typed blocks with delta and ordinal dictionary compression" /><p>Dimension fields are declared in the index mapping; the planner knows the schema up front, and grouping keys stay narrow and predictable.</p><p>In the columnar engine, each grouping label is carried as a separate vector or block. A key with 10 labels therefore requires 10 vectors to be read, hashed, compared, and retained by aggregation operators. As key width grows, so does the amount of data and bookkeeping that must move through the pipeline:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt13b25c686bbb46e0/6aa39429de2395f2e2e86d1d/7.png" alt="PromQL aggregation without packing: each grouping label hashed as a separate block into 64-byte keys" /><p>To avoid paying that per-column cost for every additional label, we introduced dimension packing. Before aggregation begins, the engine encodes the full grouping key into a single compact representation. Hash and comparison operations run on the packed key rather than on each block independently:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5a75855ef51c017b/6aa394468406d95667caa1de/8.png" alt="Dimension packing in Elasticsearch encodes PromQL grouping labels into one 16-byte key before hashing" /><p>Packing lets hashing and comparison operate on a single compact key rather than an increasing number of grouping blocks, making aggregation overhead less sensitive to key width. Because the engine is shared, ES|QL time-series queries will benefit from this optimization as well.</p><h3><strong>Building the Prometheus HTTP API response inside the pipeline</strong></h3><p>Unlike Elasticsearch's ES|QL column-oriented response format, the Prometheus response is row-oriented.  The Prometheus API returns one result row per time series, with its samples represented as timestamp-value pairs. </p><p>To support a compatible API layer, we had to regroup in the HTTP layer converting the columnar results into boxed row objects and accumulate them in map- and list-based structures until the complete Prometheus response could be produced. </p><p>We replaced this with the <code>TimeSeriesCollapse</code> compute operator. It groups rows by series and aligns samples to the query’s fixed step grid. It emits the reshaped result as ordinary columnar pages containing one row per series, with aligned multi-valued timestamp and value blocks. And it preserves compact, vectorized block representation throughout the pipeline: </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt99e2ee3137bc2fa6/6aa39496d29b4e02e81da7b6/9.png" alt="TimeSeriesCollapse operator reshapes five columnar rows into two Prometheus time series per output page" /><p>The HTTP layer can now serialize those blocks directly, avoiding the maps, lists, boxed objects, and associated allocations required by the earlier implementation.</p><h2><strong>Testing PromQL compatibility against 2,000 real queries</strong></h2><p>Prometheus <a href="https://github.com/prometheus/compliance">compliance tests</a> were our starting point.</p><p>Even though they gave us a strong baseline, they didn’t tell us how frequently individual PromQL features appear in real workloads. To complement that baseline, we built a second test corpus from over 2,000 PromQL queries collected from public repositories. </p><p>We then classified those queries by the language features and expression patterns they exercise:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf008821c47e46040/6aa394b68406d90c5bcaa1e2/10.png" alt="PromQL feature use across 2,000 real queries: aggregations 59.67%, selectors 57.55%, rate functions 45.21%" /><p>For each compatible query shape, we run the same query against Elasticsearch and Prometheus and compare the results. </p><p>In addition to that, we actively rely on <a href="https://en.wikipedia.org/wiki/Fuzzing">fuzz testing</a>, which catches issues that unit tests alone are unlikely to expose, including differences in timestamp alignment, label retention, aggregation behavior, range-vector evaluation, and response encoding.</p><h2><strong>Which PromQL functions and APIs are supported in 9.5</strong></h2><p>Since 9.4 (technical preview), PromQL support in Elasticsearch has expanded substantially. In Elasticsearch 9.5, both the<a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/promql"> <code>PROMQL</code></a> command in ES|QL and the<a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api"> Prometheus HTTP APIs</a> are generally available (GA), with more than 80% of the PromQL workflows in our real-world corpus now running without modification:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6706c16409dccfd7/6aa394d327a5312acfdcbf41/1.png" alt="PromQL compatibility in Elasticsearch rising from zero to 80% between 9.4 Tech Preview and 9.5 GA" /><p>The main additions since technical preview are:</p><p><strong>Feature</strong></p><p><strong>Example</strong></p><p><strong>Status</strong></p><p>Prometheus remote write ingest</p><p><code>POST /_prometheus/api/v1/write</code></p><p>GA in 9.5</p><p>Range queries</p><p><code>/api/v1/query_range</code></p><p>GA in 9.5</p><p>Instant queries</p><p><code>/api/v1/query</code></p><p>GA in 9.5</p><p>Metric metadata and build info</p><p><code>/api/v1/metadata</code>, <code>/api/v1/status/buildinfo</code></p><p>GA in 9.5</p><p>Native histogram functions</p><p><code>histogram_quantile</code>, <code>histogram_count</code>, <code>histogram_sum</code></p><p>GA in 9.5</p><p>Per-selector offset modifiers</p><p><code>[5m] offset 1h</code></p><p>GA in 9.5</p><p>Top-level <code>or</code> operator</p><p><code>rate(a[5m])</code> or <code>rate(b[5m])</code></p><p>GA in 9.5, up to eight operands</p><h3><strong>Prometheus remote write ingest</strong></h3><p>Elasticsearch <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/tsds-ingest-prometheus-remote-write">accepts</a> Prometheus remote write (v1) messages directly:</p><p></p><p></p><p>Snappy-compressed Protocol Buffer messages are decoded, and labels are mapped to TSDS dimensions. Metric names and values are written into the time-series index. The built-in template is dynamic, so users don’t have to predeclare every Prometheus label or metric.</p><h3><strong>Range and instant queries through the Prometheus HTTP API</strong></h3><p>Both range and instant query endpoints are <a href="https://www.elastic.co/search-labs/blog/elasticsearch-native-prometheus-api">available</a>:</p><p></p><p></p><p></p><p></p><p>Range queries return matrices evaluated over a time window, and instant queries return vectors evaluated at a single timestamp. These endpoints can be used by Kibana, Grafana, or Prometheus-compatible alerting tools, and custom dashboards.</p><h3><strong>Metric metadata and build info endpoints</strong></h3><p>Elasticsearch <a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api#promql-http-api-metadata">exposes</a> metadata about available metrics and a build-info endpoint:</p><p></p><p></p><p></p><p></p><p>The metadata endpoint returns metric types and help text, and the build-info endpoint returns the Prometheus-compatible server version. Grafana and other tools use these endpoints for feature detection and UI behavior.</p><h3><strong>Native histogram functions: histogram_quantile, count, and sum</strong></h3><p>Elasticsearch supports the <a href="https://www.elastic.co/docs/reference/query-languages/promql/functions/histogram">main PromQL operations</a> over native histograms:</p><p></p><p></p><p></p><p></p><p>Native histograms adapt their bucket layout to the data, providing useful precision across a wide value range without requiring users to configure every bucket boundary in advance. Classic histograms continue to work alongside native histograms.</p><h3><strong>Per-selector offset modifiers in PromQL</strong></h3><p>Offset modifiers shift a selector’s time window backward:</p><p></p><p></p><p>This returns the request rate from one hour earlier. Per-selector offsets are commonly used to compare current traffic, latency, or resource usage with an earlier baseline, such as the same period one week ago.</p><h3><strong>Top-level </strong><strong><code>or</code></strong><strong> operator in PromQL</strong></h3><p>Elasticsearch supports the top-level PromQL <code>or</code> operator:</p><p></p><p></p><p>In PromQL, <code>or</code> isn’t a Boolean operation. It performs a union between two sets of time series. Results from the left side are retained; a series from the right side is added only when its label set doesn’t match a series already returned by the left side. This is useful during migrations where the same logical metric may exist under an old and a new name.</p><p>The implementation follows Prometheus’s left-side precedence rules and preserves the <code>__name__</code> label. Top-level chains of up to eight operands are supported.</p><h2><strong>PromQL features not yet supported in Elasticsearch</strong></h2><p>GA doesn’t mean complete PromQL compatibility. Some less common and more complex parts of PromQL remain unsupported. These gaps now define the next phase of the work: </p><p><strong>Feature</strong></p><p><strong>Example</strong></p><p><strong>Status</strong></p><p>Advanced vector matching</p><p><code>on(instance) group_left</code></p><p>Planned</p><p>Sorting and ranking</p><p><code>topk</code>, <code>bottomk</code>, <code>limitk</code>, <code>sort</code>, <code>sort_desc</code></p><p>Planned</p><p>Label manipulation</p><p><code>label_replace</code>, <code>label_join</code></p><p>Planned</p><p>Absolute time modifier</p><p><code>@ 1710000000</code></p><p>Planned</p><p>Mixed-offset compound expressions</p><p><code>rate(...) - rate(... offset 1h)</code></p><p>Planned</p><p>Alerting and target endpoints</p><p><code>/api/v1/alerts</code>, <code>/api/v1/targets</code></p><p>Out of scope</p><h3><strong>Advanced </strong><a href="https://prometheus.io/docs/prometheus/latest/querying/operators/#group-modifiers"><strong>vector matching</strong></a><strong> with </strong><strong><code>on()</code></strong><strong> and </strong><strong><code>group_left</code></strong></h3><p>Some binary operations that require Prometheus vector matching aren’t yet part of GA.</p><p>For example, this query divides per-instance request rates by a per-instance capacity metric:</p><p></p><p></p><p>The <code>on(instance)</code> clause specifies which labels identify matching series. <code>group_left</code> permits many request-rate series to match a single per-instance capacity series, while retaining the labels from the higher-cardinality left-hand side.</p><p>These expressions are common when joining a detailed metric with metadata or a lower-cardinality capacity metric. Basic binary expressions are supported where applicable, while the remaining vector-matching forms are planned work.</p><h3><strong>Sorting and ranking: </strong><strong><code>topk</code></strong><strong>, </strong><strong><code>bottomk</code></strong><strong>, and </strong><strong><code>sort</code></strong></h3><p>Prometheus sorting and ranking functions are also not yet part of GA:</p><p></p><p></p><p>This returns the 10 services with the highest request rate. Similar queries are widely used in “top offenders” dashboards for traffic, latency, errors, and resource consumption.</p><p>The remaining functions include:</p><p></p><p></p><p></p><p></p><p></p><p></p><h3><strong>Label manipulation with label_replace and label_join</strong></h3><p>PromQL can construct or rewrite labels during query evaluation. These functions are particularly useful when dashboard variables, naming conventions, or label schemas don’t match exactly:</p><p></p><p></p><p>This creates an <code>environment</code> label from the <code>cluster</code> label.</p><p>Another common example combines existing labels into a display-oriented label:</p><p></p><p></p><p>This produces a <code>target</code> label, such as <code>payments/api-7f6d9</code>. <code>label_replace(...)</code> and <code>label_join(...)</code> aren’t yet included in GA.</p><h3><strong>Advanced time modifiers: The </strong><strong><code>@</code></strong><strong> modifier and mixed offsets</strong></h3><p>Several advanced time modifiers and expression forms remain outside the GA scope.</p><p>For example, an absolute <code>@</code> modifier evaluates a selector at a fixed Unix timestamp rather than at the query’s normal evaluation time:</p><p></p><p></p><p>This is useful for comparisons against a fixed historical point.</p><p>PromQL also permits expressions in which the two sides use different offsets:</p><p></p><p></p><p>This compares current traffic with traffic one hour earlier. Per-selector <code>offset</code> is available in GA, but not every combination of offsets and compound expressions is part of GA yet.</p><h3><strong>Prometheus API endpoints not yet implemented</strong></h3><p>In addition, the Prometheus HTTP API surface isn’t yet fully complete. Notably:</p><p>Alerting metadata through:</p><p></p><p></p><p>used by tools that inspect active alert state.</p><p>Target discovery through:</p><p></p><p></p><p>used to inspect scrape targets, health, and labels.</p><p>These endpoints concern Prometheus server and scrape-target state rather than querying metrics stored in Elasticsearch.</p><p>For the full list of limitations, see the <a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-limitations#promql-limitations-form-post">PromQL limitations</a> page. </p><h2><strong>Try PromQL in Elasticsearch 9.5</strong></h2><p>To query Prometheus metrics in Elasticsearch 9.5 or Serverless, see the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/promql">PromQL documentation</a> and the <a href="https://www.elastic.co/docs/reference/query-languages/promql/promql-http-api">Prometheus HTTP API reference</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/promql-elasticsearch-compute-engine</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/promql-elasticsearch-compute-engine</guid>
    <category><![CDATA[Query Languages]]></category>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Sergey Sidorov,Felix Barnsteiner]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf690e82ea51bbeec/6aa37da41ade6445cc42eded/unnamed.png" length="0" type="image/png"/>
    <pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Don't leave metrics on the table: query them with the ES|QL TS command]]></title>
    <description><![CDATA[Recalibrate your mental model for time series queries: learn why FROM can produce inaccurate results for metrics, how TS fixes that, and when to use each command.]]></description>
    <content:encoded><![CDATA[<p>If you use ES|QL for logs and traces, <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/from"><code>FROM</code></a> is probably second nature, but on metrics it can return numerically wrong answers. A query like <code>FROM metrics-* | STATS SUM(request_count)</code> adds up cumulative counter values across every sample on every host. The result grows without bound and isn't a rate, a count, or anything else useful. <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts"><code>TS</code></a> fixes that by grouping samples into time series first, then exposing functions like <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-rate"><code>RATE</code></a>, <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-avg_over_time"><code>AVG_OVER_TIME</code></a>, and <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-last_over_time"><code>LAST_OVER_TIME</code></a> that operate per series.</p><p>For a high-level tour of metrics analytics across ES|QL and Discover, see <a href="https://www.elastic.co/observability-labs/blog/metrics-explore-analyze-with-esql-discover">Explore and Analyze Metrics with Ease in Elastic Observability</a>. This post zooms in on the mechanics.</p><p>Here is the mental model in five bullets:</p><ul><li><p><code>FROM</code> treats every document as an independent row. That is right for events, but metric aggregations often need the time series that each row belongs to.</p></li><li><p><code>TS</code> adds that time series context: it groups and aggregates data points by time series before any other aggregation runs, and enables functions like <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-rate"><code>RATE</code></a>, <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-avg_over_time"><code>AVG_OVER_TIME</code></a>, and <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-last_over_time"><code>LAST_OVER_TIME</code></a>.</p></li><li><p>A <code>TS | STATS</code> query normally has two aggregation phases. The inner phase reduces samples inside each time series; the outer phase groups and combines those per-series results.</p></li><li><p>The default inner aggregation is <code>LAST_OVER_TIME</code>, which is why <code>TS metrics | STATS AVG(cpu_usage)</code> and <code>FROM metrics | STATS AVG(cpu_usage)</code> can return different numbers.</p></li><li><p>Use <code>TS</code> to query a time series data stream (<a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/time-series-data-stream-tsds">TSDS</a>). Use <code>FROM</code> for events and raw document inspection.</p></li></ul><h2>What is a time series, really?</h2><p>A time series is a sequence of <code>(timestamp, value)</code> data points identified by the metric name and a unique set of dimension values.</p><p>For example, <code>request_count</code> reported every 30 seconds by host <code>h1</code> in data center <code>dc1</code> is one time series. The same metric on host <code>h2</code> in <code>dc1</code> is a different time series.</p><p>In a time series data stream, every metric document carries an internal <code>_tsid</code> field that uniquely identifies a time series. Samples that share a <code>_tsid</code> belong to the same time series and are stored sequentially, sorted by timestamp.</p><p>That storage layout enables efficient per-series aggregations. It also explains why <code>TS</code> only works on time series data streams. Other index modes have no notion of a time series, so the per-series operations <code>TS</code> relies on have no such identifier to attach to. <code>FROM</code> does not support those operations, which is what the next section is about.</p><h2>Why FROM leaves metrics on the table</h2><p>Consider a counter named <code>request_count</code> collected every 30 seconds from three hosts.</p><p>A counter is a cumulative metric: each sample is the running total since the process started reporting it. For <code>request_count</code>, a value of <code>1,000</code> means "this time series has observed 1,000 requests so far", not "1,000 requests happened since the previous sample". Counters reset to zero on process restart, so a sample of <code>4</code> right after <code>1,004</code> is a fresh count, not negative traffic. The ES|QL <code>RATE</code> function computes the per-second change within a time series and handles resets without glitches.</p><p>You want to calculate the total request rate across all hosts, bucketed by 5 minutes.</p><p>If you are used to writing ES|QL over event data, you might start with this query:</p><p>The chart it produces looks plausible at first: a line that goes up over time. But the number on the y-axis is the sum of every cumulative counter value reported in the bucket. Each host contributes its own running total, repeatedly, once per sample. Because the query uses <code>SUM</code> on those cumulative values, the result is not a rate, it is not the number of requests in the bucket, and it grows without bound even if the application stops receiving requests.</p><p><code>request_count</code> is a monotonically increasing counter, so its raw values represent "how many requests have ever happened on this host", not how many happened in the bucket. The right computation is "how much did this counter increase per second on each host, then sum across hosts." <code>FROM</code> cannot express that operation directly. It can group rows by fields, but it has no built-in notion of "the same time series over time" and no way to ask for the change of a counter within each time series. It also cannot use sliding-window time series functions such as <code>RATE(request_count, 5m)</code>, which we will come back to below.</p><p><code>TS</code> was introduced for this purpose, providing a succinct syntax to express time series aggregations:</p><p><code>RATE(request_count)</code> runs per time series and produces a per-second rate that handles counter resets correctly. <code>SUM</code> then adds those rates across hosts.</p><h2>Two aggregation phases: inner and outer</h2><p>Every <code>TS | STATS</code> query has two distinct aggregation phases.</p><p>Let's make that concrete with a query that calculates the request rate per data center:</p><p>The diagram below shows how <code>TS</code> evaluates this query. It first reduces samples inside each time series, then groups and combines those per-series values into one result per <code>datacenter</code> and time bucket.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte5b1f3d827f95ad1/6a1706f0961e6910afc4ce92/1d3d765ab8e27e539ac6bedf3e7444632a4bbe7e-3050x900.png" alt="Inner and outer aggregation phases of a TS|STATS query" /><p>The phases are:</p><p><strong>Inner (within a time series).</strong> Runs separately for each time series. It collapses many <code>(timestamp, value)</code> data points within a bucket into a single value per time series per bucket by applying the inner aggregation function, such as <code>RATE</code> in the example above. Functions: <code>RATE</code>, <code>AVG_OVER_TIME</code>, <code>MAX_OVER_TIME</code>, <code>LAST_OVER_TIME</code>, <code>STDDEV_OVER_TIME</code>, and so on. The full list is on the <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions">time series aggregation functions</a> page.</p><p><strong>Outer (across time series, the "grouping" phase).</strong> Combines the per-series values into a single value per group per bucket. Functions: <code>SUM</code>, <code>AVG</code>, <code>MAX</code>, <code>MIN</code>, percentiles, and the rest of the <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/aggregation-functions">regular ES|QL aggregates</a>.</p><p>In <code>SUM(RATE(request_count)) BY datacenter, TBUCKET(5m)</code>:</p><ul><li><p><code>RATE(request_count)</code> is the inner aggregation. It runs per time series.</p></li><li><p><code>SUM(...)</code> is the outer aggregation. It combines time series within the same <code>datacenter</code> and bucket.</p></li><li><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/grouping-functions#esql-tbucket"><code>TBUCKET(5m)</code></a> defines the bucket boundaries (equivalent to <code>BUCKET(@timestamp, 5m)</code>).</p></li></ul><p>The outer aggregation is optional. If you only need the per-time-series result, use the time series aggregation function directly:</p><p>That query keeps the per-series rate for each bucket instead of wrapping it in <code>SUM</code>, <code>AVG</code>, or another aggregate across time series.</p><h2>The default inner aggregation: LAST_OVER_TIME</h2><p><code>TS</code> has to reduce raw samples inside each time series before it can run the outer aggregation. That means every metric field in a <code>TS | STATS</code> aggregation needs an inner aggregation, even when the query does not spell one out.</p><p>Consider a metric named <code>cpu_usage</code>. It is a gauge: a metric that captures a value at a point in time and can move up and down freely. A sample of <code>0.42</code> means "this host is at 42% CPU at this time". For a gauge, the natural "value in this bucket" is the most recent sample.</p><p>That is what ES|QL fills in for you. If you write <code>TS metrics | STATS AVG(cpu_usage) BY host.name, TBUCKET(5m)</code>, the implicit inner aggregation is <code>LAST_OVER_TIME(cpu_usage)</code> and the query is equivalent to:</p><p>For each time series, <code>LAST_OVER_TIME</code> picks the latest sample in the bucket. Then <code>AVG</code> averages across time series.</p><p>It is also why the same-looking query against <code>FROM</code> and <code>TS</code> can return different numbers. <code>FROM</code> averages every individual document. <code>TS</code> averages one value per time series per bucket. If your hosts publish at slightly different rates, those averages diverge. For example, in a five-minute bucket, a host that publishes every second contributes 300 documents while a host that publishes every two minutes contributes only two or three. With <code>FROM | STATS AVG(cpu_usage)</code>, the chatty host dominates the average. With <code>TS</code>, each time series is reduced to one bucket value first, so the outer average gives each host one value to contribute.</p><p>If you want the average value during the bucket instead of the latest value, make the inner aggregation explicit:</p><p><code>AVG_OVER_TIME</code> averages all CPU utilization samples within each time series. The outer <code>AVG</code> then averages those per-series values across matching hosts. That makes the result sample-weighted within each time series, then equally weighted across time series. Use this when you care about how the value behaved during the bucket, not just where it ended up.</p><p>The same rule applies to peaks and troughs. For a peak CPU chart, use <code>MAX(MAX_OVER_TIME(cpu_usage))</code>, not just <code>MAX(cpu_usage)</code>. The inner <code>MAX_OVER_TIME</code> finds the peak within each time series; the outer <code>MAX</code> finds the peak across matching time series.</p><p>Counters work the other way around. Their sample value is a running total, so the latest sample on its own is rarely meaningful. For a counter, the inner aggregation you almost always want is <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions#esql-rate"><code>RATE</code></a> for a per-second rate, or <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions"><code>INCREASE</code></a> for the total change in the bucket. Falling back on the default <code>LAST_OVER_TIME</code> gives you the most recent cumulative value, which is the trap the FROM query in the previous section walked into.</p><p>Pick the inner function deliberately. The outer function is the easy part.</p><h2>When to use TS, when to use FROM</h2><p>A practical rule of thumb:</p><ul><li><p>Use <code>TS</code> for metric aggregations against a <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/time-series-data-stream-tsds">time series data stream</a>. It is the source command designed for that data, and it applies per-series semantics by default.</p></li><li><p>Use <code>FROM</code> for events: logs, traces, audit records, transactions. Each row is independent. There is no time series context.</p></li></ul><p><code>FROM</code> still works on TSDS indices and is occasionally useful, for example when you want to inspect raw metric documents without per-series grouping. For dashboards, alerts, and any kind of charting, <code>TS</code> is the right default.</p><p>If you first need to discover which metrics or time series exist in the data, use <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/metrics-info"><code>METRICS_INFO</code></a> or <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts-info"><code>TS_INFO</code></a> after <code>TS</code> and before <code>STATS</code>. See <a href="https://www.elastic.co/search-labs/blog/esql-metrics-info-ts-info-time-series-catalog">ES|QL METRICS_INFO and TS_INFO: Catalog your time series data</a> for a deeper walkthrough.</p><h2>Post-process TS results with ES|QL</h2><p>The first <code>STATS</code> command is the boundary between time series processing and regular ES|QL processing. Before that first <code>STATS</code>, <code>TS</code> needs to keep the data grouped by <code>_tsid</code>, so commands that change row order or shape are not allowed. After that first <code>STATS</code>, the output is a regular ES|QL table. You can sort it, limit it, join lookup data, enrich it, or compute derived columns.</p><p>For example, this query calculates average CPU per host and bucket, finds the maximum bucketed average for each host, and returns the ratio:</p><h2>Sliding windows for the inner aggregation</h2><p>Time series aggregation functions accept a second argument: the window size for the inner phase.</p><p>This computes the rate over a 5-minute sliding window, but reports a value every minute. It is useful when you want a smoother chart at fine bucket sizes.</p><p>The window is the ES|QL counterpart to a PromQL <a href="https://prometheus.io/docs/prometheus/latest/querying/basics/#range-vector-selectors">range vector selector</a>: <code>RATE(app.requests, 5m)</code> serves the same purpose as <code>rate(app_requests[5m])</code>.</p><h2>Gotchas worth knowing</h2><p>A few things in <code>TS</code> can seem surprising, especially when coming from the events-based <code>FROM</code> mental model. None of these are bugs; most are direct consequences of the per-series model. Here is what to watch for.</p><p><strong><code>COUNT(*)</code></strong> <strong>is rejected.</strong> Say you want to know how many samples were collected per service in each bucket. The instinct from <code>FROM</code> is <code>COUNT(*)</code>, but <code>TS</code> rejects it: there is no plain "row" once data is grouped by time series, so a row count has no defined meaning. Pick what you actually want to count:</p><ul><li><p>Number of samples per service: <code>STATS samples = SUM(COUNT_OVER_TIME(cpu_usage)) BY service.name, TBUCKET(5m)</code>. The inner <code>COUNT_OVER_TIME</code> counts samples per time series; the outer <code>SUM</code> adds them across the time series in the group.</p></li><li><p>Number of distinct hosts reporting per service: <code>STATS hosts = COUNT_DISTINCT(host.name) BY service.name, TBUCKET(5m)</code>. This counts unique label values across time series.</p></li></ul><p><strong>You cannot sort, limit, lookup join, or enrich before</strong> <strong><code>STATS</code></strong><strong>.</strong> <code>TS metrics | SORT @timestamp | STATS ...</code> will fail. The grouping by <code>_tsid</code> must happen first, before anything else can run. Filter with <code>WHERE</code> if you need to narrow the scope. After the first <code>STATS</code>, the output is regular ES|QL and you can pipe it through any command, as shown in the previous section.</p><p><strong>Gauge vs counter mapping.</strong> Time series functions are sensitive to the metric type set in the field mapping. <code>RATE</code> only works on counters; <code>*_OVER_TIME</code> functions are intended for gauges. If you build TSDS mappings by hand, pay special attention to this part.</p><p>This can be a source of friction for Prometheus users. Prometheus metric type metadata is not always available in the data Elasticsearch receives, so the metric type may have to be inferred from naming conventions (<code>_total</code> for counters, and so on). Those heuristics are imperfect, and a misclassified metric is rejected by the function that should accept it. The deeper mechanics, including how Prometheus Remote Write maps metric types into TSDS, are covered in <a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">How Prometheus Remote Write Ingestion Works in Elasticsearch</a>.</p><p>Explicit converter functions (gauge-to-counter and counter-to-gauge) are on the roadmap to make these cases easier to recover from at query time.</p><p><strong>Kibana charts go empty when you zoom in too far.</strong> In Kibana, <code>TBUCKET</code> adapts to the date picker, so zooming in shrinks the bucket size. When the bucket size drops below the data's collection interval, every other bucket has no sample, <code>RATE</code> and the rest return null, and the chart silently goes blank. Elastic is evaluating mitigations such as a runtime warning when the bucket size is too small, a configurable minimum bucket size, or automatic widening of the window or bucket size.</p><h2>Wrap up</h2><p>For metric queries, start with <code>TS</code> unless you specifically need raw documents. Then choose the inner aggregation based on what the value should mean inside each time series: <code>RATE</code> for counters, <code>LAST_OVER_TIME</code> for current gauge values, and explicit <code>*_OVER_TIME</code> functions for peaks, averages, minimum values, or distributions.</p><p>Once the per-series value is right, the outer aggregation is the familiar part: group and reduce those time series into the chart, alert, or table you need.</p><p>For the full reference, see the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts"><code>TS</code></a> <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/ts">command docs</a> and the list of <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/time-series-aggregation-functions">time series aggregation functions</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-ts-command-querying-metrics</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-ts-command-querying-metrics</guid>
    <category><![CDATA[ES|QL]]></category>
    <dc:creator><![CDATA[Felix Barnsteiner]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt81635cd2cc7703b3/6a1706ef1949f7a977e7a95c/e2eb1ba006612a352f1317c1621e4ebc5b2a12b6-1376x768.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Bringing Fire to Elasticsearch: Adding Native Prometheus API Support]]></title>
    <description><![CDATA[Query Elasticsearch directly from Prometheus-compatible clients via native PromQL, discovery, and metadata endpoints. Send data to Elasticsearch with Prometheus Remote Write.]]></description>
    <content:encoded><![CDATA[<p>Point any Prometheus-compatible client at Elasticsearch and run PromQL directly against your existing metrics. Elasticsearch is adding native Prometheus query, discovery, and metadata endpoints as a tech preview that work over metrics ingested through Prometheus Remote Write, OpenTelemetry, or the Bulk API. The API runs on top of Elasticsearch's time series data streams (TSDS), so there's no separate Prometheus-specific storage layer to operate.</p><p>This post explains how the query, discovery, and metadata endpoints build on the earlier ingest and query work to form that API surface. Companion posts go deeper on individual pieces:</p><ul><li><p><a href="https://www.elastic.co/observability-labs/blog/elasticsearch-supports-promql">Native PromQL support in ES|QL</a> covers how PromQL queries are translated into ES|QL execution plans.</p></li><li><p><a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch">Ship Prometheus Metrics to Elasticsearch with Remote Write</a> covers ingestion setup.</p></li><li><p><a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">How Prometheus Remote Write Ingestion Works in Elasticsearch</a> covers the remote write internals.</p></li></ul><p>This is still a work in progress. The sections below call out what is supported today and which parts are still evolving.</p><h2>The API surface</h2><p>Today, the Prometheus-compatible API surface falls into three groups.</p><h3>Query endpoints</h3><p>The query endpoints let Prometheus-compatible clients evaluate PromQL expressions:</p><ul><li><p><code>GET /_prometheus/api/v1/query_range</code> evaluates a PromQL expression over a time window (matrix results).</p></li><li><p><code>GET /_prometheus/api/v1/query</code> evaluates at a single point in time (vector results). Currently implemented as a short range query that returns the last sample.</p></li></ul><p>Only GET is supported for query endpoints today. Some clients default to POST, so you may need to configure them to use GET. The Prometheus POST convention uses <code>application/x-www-form-urlencoded</code> bodies, which Elasticsearch's HTTP layer rejects as a CSRF safeguard before the request ever reaches the handler.</p><p>For the full PromQL coverage status, see the <a href="https://www.elastic.co/observability-labs/blog/elasticsearch-supports-promql">companion post on PromQL in ES|QL</a>.</p><h3>Metadata endpoints</h3><p>The metadata endpoints serve the discovery information that clients need for autocomplete, variable dropdowns, and metric browsing.</p><p>The series, labels, and label values endpoints all accept <code>match[]</code> selectors and a time range (<code>start</code>/<code>end</code>). The <code>match[]</code> parameter takes a Prometheus series selector like <code>http_requests_total{job="api"}</code> and restricts the response to time series that match. This keeps responses fast and relevant on clusters with large numbers of metrics. For example:</p>GET /_prometheus/api/v1/series?match[]=http_requests_total{job="api"}GET /_prometheus/api/v1/labels?match[]=http_requests_totalGET /_prometheus/api/v1/label/instance/values?match[]=http_requests_total{job="api"}<p>The first returns all series for <code>http_requests_total</code> where <code>job="api"</code>, with their full label sets. The second returns only the label names that exist on <code>http_requests_total</code> series. The third returns only the <code>instance</code> values that appear on matching series.</p><p><code>GET /_prometheus/api/v1/metadata</code> is different: it returns type and unit for each metric, optionally filtered by name via a <code>metric</code> parameter.</p>GET /_prometheus/api/v1/metadata?metric=http_requests_total<p>It does not accept <code>match[]</code> selectors or a time range. In Prometheus, metadata is collected from active scrape targets (the <code>HELP</code>, <code>TYPE</code>, and <code>UNIT</code> lines they expose), so the response does not involve a data scan. Elasticsearch does not have a dedicated metadata store like that, so the current implementation discovers metric metadata by visiting time series data from the last 24 hours. This keeps the query fast without requiring a full index scan. That 24-hour lookback is fixed today: the Prometheus metadata API does not expose <code>start</code> or <code>end</code> parameters that Elasticsearch could use to make it user-adjustable.</p><p>How the metadata endpoints work under the hood, including the <code>TS_INFO</code> and <code>METRICS_INFO</code> commands that power them, is covered <a href="https://www.elastic.co/search-labs/blog//elasticsearch-native-prometheus-api#ts-info-and-metrics-info">below</a>.</p><h3>Index pre-filtering</h3><p>All query and metadata endpoints accept an optional <code>{index}</code> path segment after <code>/_prometheus/</code>:</p>GET /_prometheus/metrics-prod-*/api/v1/query_range?query=up&amp;start=...&amp;end=...<p>This restricts which Elasticsearch indices the query runs against before any expression evaluation begins. On clusters with many data streams across teams or environments, this avoids scanning unrelated indices and can significantly reduce query latency. You can configure separate data sources per index pattern to give teams scoped access to their own metrics.</p><h3>A note about Remote Write</h3><p>For ingestion, Elasticsearch also exposes the standard Prometheus Remote Write endpoint:</p><ul><li><p><code>POST /_prometheus/api/v1/write</code> ingests time series via the Prometheus Remote Write v1 protocol. v2 is not yet supported.</p></li></ul><p>Remote Write writes into Elasticsearch's existing time series data streams (TSDS), not a separate Prometheus-specific storage layer. Prometheus labels become TSDS dimensions, and metric names become fields in the index mapping. The <a href="https://www.elastic.co/observability-labs/blog/prometheus-remote-write-elasticsearch-architecture">remote write architecture post</a> covers the full mapping in detail, including how metric types are inferred and how labels are stored with a <code>labels.</code> prefix.</p><h3>How it works</h3><p>Under the hood, all endpoints work the same way: parse the incoming HTTP parameters, build an ES|QL query plan, execute it against time series data streams, and convert the columnar result back into the JSON format Prometheus clients expect.</p><h2>TS_INFO and METRICS_INFO</h2><p>The metadata endpoints need to answer questions like "what labels exist?" or "what metric types are defined?" across potentially millions of time series, without scanning every data point.</p><p>Internally, the Prometheus metadata endpoints answer those questions by building ES|QL plans around two new processing commands: <code>METRICS_INFO</code> and <code>TS_INFO</code>. You do not need to use these commands directly to use the Prometheus API, but they are the core execution primitives behind the metadata responses. Both work by visiting only one document per time series to extract its metadata, rather than scanning all samples. This means their cost scales with the number of distinct time series, not the number of data points.</p><p><code>METRICS_INFO</code> returns one row per distinct metric with its name, type, unit, and associated dimension fields. <code>TS_INFO</code> is more granular: one row per (metric, time series) combination, including the actual dimension values as a JSON object.</p><p>A dedicated blog post on <code>TS_INFO</code> and <code>METRICS_INFO</code> is coming soon, covering the two-phase execution model, how they scale, and how to use them directly in ES|QL queries beyond the Prometheus API.</p><h3>How the metadata endpoints use them</h3><p>Each metadata endpoint constructs an ES|QL plan with one of these commands at its core.</p><p><code>/api/v1/labels</code> and <code>/api/v1/series</code> use <code>TS_INFO</code>, since they need per-time-series detail (which labels exist, which dimension values identify each series). <code>/api/v1/metadata</code> and <code>/api/v1/label/__name__/values</code> use <code>METRICS_INFO</code>, since they only need per-metric information (metric names, types, units).</p><p><code>/api/v1/label/{name}/values</code> for regular labels (anything other than <code>__name__</code>) does not use either command. Regular labels like <code>job</code> or <code>instance</code> are actual dimension fields in the index, so the endpoint can query them directly with a group-by aggregation. When <code>match[]</code> selectors are provided, they are translated into a <code>WHERE</code> clause that filters the time series before the aggregation runs.</p><p>The <code>__name__</code> label needs a different strategy because it is not always present as a dimension field. Prometheus Remote Write does store <code>labels.__name__</code>, but metrics ingested through other paths (OpenTelemetry, the bulk API) do not have it. The metric name is encoded in the field name itself (e.g., <code>metrics.http_requests_total</code>). You could look at the index mappings to enumerate field names, but mappings alone do not tell you which metric has which dimensions, and they cannot be filtered by label values from a <code>match[]</code> selector. <code>METRICS_INFO</code> can do both: it enumerates metric names across indices while respecting upstream <code>WHERE</code> filters.</p><p>In all cases, the API layer handles the translation back to Prometheus conventions: stripping the <code>labels.</code> and <code>metrics.</code> storage prefixes and synthesizing <code>__name__</code> for non-Prometheus metrics that lack it.</p><h2>In conclusion</h2><p>The result: any Prometheus-compatible client can query and explore Elasticsearch metrics through endpoints it already understands. Remote Write metrics, OpenTelemetry metrics, and metrics indexed through other paths all show up through the same API, backed by the same TSDS indices.</p><p>All the Prometheus APIs mentioned here are available as tech preview in Elasticsearch Serverless today. For self-managed clusters and Elastic Cloud Hosted deployments, available as tech preview in Elasticsearch 9.4, with the exception of <code>GET /_prometheus/api/v1/metadata</code>. To experiment locally, use <a href="https://www.elastic.co/docs/deploy-manage/deploy/self-managed/local-development-installation-quickstart">start-local</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-native-prometheus-api</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-native-prometheus-api</guid>
    <category><![CDATA[Integrations]]></category>
    <dc:creator><![CDATA[Felix Barnsteiner]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt12b4e100d5bbb7f0/6a16f7a22b835ff747f4afdd/c7b333bd73e8a1f4e18486b2d692ba742788dcfd-1376x768.jpg" length="0" type="image/jpeg"/>
    <pubDate>Mon, 11 May 2026 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>