<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Valeriy Khakhutskyy - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Valeriy Khakhutskyy - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/search-labs/author/valeriy-khakhutskyy</link>
    </image>
    <link>https://www.elastic.co/search-labs/author/valeriy-khakhutskyy</link>
    <atom:link href="https://www.elastic.co/search-labs/rss/author/valeriy-khakhutskyy.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Mon, 28 Sep 2026 20:49:40 GMT</lastBuildDate>
  <item>
    <title><![CDATA[Is your ML job's datafeed losing a race it cannot win?]]></title>
    <description><![CDATA[Learn how switching from scroll-based to aggregation-based datafeeds optimizes machine learning jobs for large-scale deployments.]]></description>
    <content:encoded><![CDATA[<p>On almost every large Elastic deployment I’ve worked with, there’s an Elastic Security or Elastic Observability anomaly detection (AD) job that looks healthy but is perpetually behind. Six hours behind. Twelve. And the gap never closes.</p><p>The datafeed isn’t broken. It’s doing exactly what it was built to do: reading every raw document, across every shard, every run. On a large cluster with cross-cluster search (CCS) and a broad index pattern, like <code>logs-*</code>, that means scanning billions of documents per bucket. There’s no hardware that makes that sustainable. The datafeed will always be chasing live data and never reaching it.</p><p>The fix is to switch from the default <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-scroll"><strong>scroll-based</strong></a> datafeed configuration to an <a href="https://www.elastic.co/docs/explore-analyze/machine-learning/anomaly-detection/ml-configuring-aggregation"><strong>aggregation-based</strong></a><a href="https://www.elastic.co/docs/explore-analyze/machine-learning/anomaly-detection/ml-configuring-aggregation"> datafeed configuration</a>: Let the data nodes summarize locally, and ship only compact bucket results to the ML node. Same detections, a fraction of the load. The speedup can be dramatic. More than you might expect. The numbers are in the next section. The explanation for <em>why</em> the gap is so large is at the end of the post, for those who want to understand the mechanics.</p><p>One catch worth knowing now: Switching requires creating a new job. The old model doesn’t transfer; weeks of learned baseline are lost. <strong>The right time to make this switch is before the job has been running for months, not after.</strong> That’s the main reason to read this before you deploy.</p><h2><strong>How much faster? Scroll vs. aggregation datafeeds for ML jobs</strong></h2><p>We ran the same job two ways on production data: first scroll-based, and then aggregation-based. The job covered 13 months of history, monitoring 836,000 log events per hour in 15-minute buckets across multiple clusters.</p><p>Training on historical data with scroll-based configuration: <strong>five days of wall-clock time</strong>, 7.9 million sequential requests, and 3.5 TB transferred; with aggregations: <strong>2.3 minutes</strong>, 23 requests, and 34 MB (a 3,374× speedup). Think of it this way: If you start the scroll backfill at 9 a.m. Monday, it will finish Saturday morning. The aggregation version is done by 9:02 a.m.</p><p>On live data, the difference is less dramatic but still meaningful: around <strong>20×</strong> fewer requests per tick. That adds up quickly when the datafeed runs every few minutes around the clock.</p><h2><strong>Before you start</strong></h2><p>Three things worth knowing before diving into the configuration.</p><p><strong>This isn't wizard territory.</strong> The standard Kibana job wizards (Single Metric, Multi-Metric, Population) don't expose aggregation configuration. To create an aggregation-based job, you need either the Elasticsearch API or Kibana's Advanced Job Wizard, with JSON edited by hand. The worked example below shows the most practical path: Configure the job in the Multi-Metric Wizard, and then click <strong>Convert to advanced job</strong> before creating it. That gets you a prefilled JSON starting point instead of a blank editor.</p><p><strong>The configuration is unforgiving and mostly silent about it.</strong> There's no schema validation that catches a misnamed aggregation key or a <code>fixed_interval</code> that doesn't match <code>bucket_span</code>. The job will run, anomalies will fire, and nothing will indicate that the results are based on the wrong data. This is why the five-step pattern exists and why the <strong>Preview </strong>tab is worth using every time: Catching a misconfiguration before the job trains is a 30-second check; catching it a week later is a much worse afternoon.</p><p><strong>The Single Metric Viewer has a known limitation with aggregated jobs.</strong> That viewer reconstructs the "actual" data curve by re-querying the index, but it can't reproduce an arbitrary, user-defined aggregation, so the actual-value line is typically missing or approximate. The Anomaly Explorer is unaffected: Anomaly scores, swim lanes, and influencer attribution all work normally. Just don't rely on the Single Metric Viewer's chart for visual validation of what the model saw.</p><h2><strong>What we can and can’t aggregate</strong></h2><p>Almost every <a href="https://www.elastic.co/docs/reference/machine-learning/machine-learning-functions">ML function</a> works with aggregated datafeeds, but the right aggregation pattern depends on the function.</p><p>Function</p><p>Pattern</p><p>`count`, `mean`, `high_mean`, `low_mean`, `sum`, `max`, `min`</p><p>Standard: `date_histogram` → `terms` → metric aggregation</p><p>`time_of_day`, `time_of_week`</p><p>Minimal: plain `date_histogram`, no `terms` or metric needed</p><p>`rare`, `freq_rare`, `info_content`</p><p>Composite: top-level composite with `date_histogram` as a source</p><p>`categorization`</p><p>`terms` on the `.keyword` subfield of the categorization field</p><p>`lat_long`, `varp`</p><p>Scroll only</p><p><code>lat_long</code> and <code>varp</code> are the genuine exceptions. If you want to use these detectors, you are required to use the scroll-based datafeed configuration.</p><p>The five-step pattern in the next section covers the standard case. We’ll walk through the remaining patterns at the end of the post.</p><h2><strong>The standard five-step pattern: Scroll-based to aggregation datafeed</strong></h2><p>Converting any scroll-based job to an aggregation-based datafeed follows the same five steps. Once you understand the pattern, applying it to any compatible job takes about 10 minutes.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4144dac3e60cccbf/6a16fa0514b2704508e3c412/77cd16165133374a04dbcf71210ea8d36f66b54f-1999x924.png" alt="Flowchart illustrating how to configure Elasticsearch ML datafeed aggregations, showing steps for summary fields, bucket topology, timestamp handling, field mapping, and detector metrics." /><p><strong>Step 1: Add </strong><strong><code>summary_count_field_name: </code></strong><a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/mapping-doc-count-field"><strong><code>"doc_count"</code></strong></a><strong> to the analysis config.</strong> This tells the ML engine that incoming data is pre-summarized. Without it, the engine treats each aggregated bucket as a single raw document and produces wrong anomaly scores.</p><p><strong>Step 2: Choose the bucket wrapper topology.</strong> For most functions (<code>count</code>, <code>mean</code>, <code>sum</code>, <code>max</code>, <code>min</code>, <code>varp</code>, <code>time_of_day</code>, <code>time_of_week</code>, and <code>categorization</code>) use a <a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-bucket-datehistogram-aggregation"><code>date_histogram</code></a> at the top level whose <code>fixed_interval</code> matches your <code>bucket_span</code> exactly to ensure accurate analysis. For <code>rare</code>, <code>freq_rare</code>, and <code>info_content</code>, use a <a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-bucket-composite-aggregation">composite</a> at the top level with a <code>date_histogram</code> as one of its sources. This routes the datafeed to the composite extractor, which paginates through all field-value combinations rather than truncating to a top-N.</p><p><strong>Step 3: Add a </strong><a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-metrics-max-aggregation"><strong><code>max</code></strong></a><strong> aggregation on </strong><strong><code>@timestamp</code></strong><strong>.</strong> The ML engine needs this to determine the precise end time of each bucket. In the standard topology (Step 2, <code>date_histogram</code> outer), it goes inside the histogram’s <code>aggregations</code>. In the composite topology, it sits as a sibling of the <code>composite</code> aggregation.</p><p><strong>Step 4: Map each analysis field to a </strong><a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-bucket-terms-aggregation"><strong><code>terms</code></strong></a><u><strong> aggregation</strong></u>, named exactly after the corresponding field in the analysis config. One categorical field → a single nested <code>terms</code>. Two or more categorical fields → a <code>composite</code> aggregation nested inside the <code>date_histogram</code>, with one <code>terms</code> source per field. For categorization jobs, use a <code>terms</code> aggregation on the <code>.keyword</code> subfield of the <code>categorization_field_name</code>. The naming rule is strict: The aggregation key must exactly match the field name in the analysis config; the ML engine uses the aggregation name, not the <code>field</code> parameter, to look up values. A mismatch produces silently wrong results; no error, just a job that appears to run while missing everything meaningful.</p><p><strong>Step 5: Map each detector’s metric field</strong> to its Elasticsearch aggregation equivalent:</p><p>ML function</p><p>Elasticsearch aggregation</p><p>`mean` / `high_mean` / `low_mean`</p><p>`avg`</p><p>`sum`</p><p>`sum`</p><p>`max`</p><p>`max`</p><p>`min`</p><p>`min`</p><p>For <code>count</code>, <code>rare</code>, <code>freq_rare</code>, <code>info_content</code>, <code>time_of_day</code>, <code>time_of_week</code>, and categorization jobs, the ML engine works from <code>doc_count</code> alone; no metric aggregation is needed, and this step can be skipped.</p><h2><strong>Step-by-step example: Building an aggregation-based ML job in Kibana</strong></h2><p>Let’s build this end to end using Kibana’s sample web logs. If you haven’t loaded them yet, go to the Kibana home page and click <strong>Integrations → Sample data → Sample web logs → Add data</strong>. This gives us a data view called <code>Kibana Sample Data Logs</code> and an index called <code>kibana_sample_data_logs</code> with fields including <code>@timestamp</code>, <code>bytes</code> (response size), and <code>geo.dest</code> (destination country).</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbd1e470b4b641b8a/6a16fa072d2f505cd7c3c0e4/75b692b9f38017cd7e4e221d2e89a14f75d3b9dc-1999x1905.png" alt="Elastic “Add data” page showing sample datasets, including ecommerce orders, flight data, and web logs, with the web logs option highlighted." /><p>We’ll build a job that detects unusually large response sizes: <code>high_mean of bytes</code>, partitioned by destination country (<code>geo.dest</code>), with a 1-hour bucket span.</p><h3><strong>Creating the job with the Multi-Metric Wizard</strong></h3><p>This is how most jobs get created in practice. Navigate to <strong>Machine Learning → Anomaly Detection → Manage Jobs → Create job</strong>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb47f8491b3a60622/6a16fa09acf088a614be98f6/b5bb2f2770a76fd535db22b97fc4f72471c43ca7-1999x587.png" alt="Kibana interface showing the “Create job” step for anomaly detection, with a panel listing available data views and the “Kibana Sample Data Logs” option selected." /><p>Select the “Kibana Sample Data Logs” data view, and set the time range to cover the full sample dataset. On the job type screen, choose <strong>Multi-metric</strong>.</p><p>In the Multi-Metric Wizard, configure the detector:</p><ul><li><p><strong>High mean</strong> of <code>bytes</code>.</p></li><li><p><strong>Split data by</strong> <code>geo.dest</code>.</p></li><li><p><strong>Bucket span:</strong> <code>1h</code>.</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcd0c896c1510795c/6a16fa0bab7f084ea8db9c62/b3055c91c881e4011521ba0c17cc36c6138595ee-1999x1540.png" alt="Kibana anomaly detection job summary showing a multi‑metric chart split by geographic destination and a configuration panel listing job ID, bucket span, split field, influencers, memory limit, and time range." /><p>Give the job an ID, and leave everything else at its defaults, but <strong>don’t click Create yet</strong>. On this last configuration step, click on <strong>Preview JSON</strong> and look at the datafeed section. What you’ll see is a plain scroll-based datafeed with no aggregations, just an index pattern and a <code>match_all</code> query.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfc768a5bb02a3a12/6a16fa0d92262a59d61cc0a5/32a1650958525b03a6052d480152933341acdd41-1999x1392.png" alt="Side‑by‑side JSON showing an Elasticsearch ML job configuration and its matching datafeed configuration, including detectors, influencers, index selection, query, and runtime mappings." /><p>This is the default every wizard produces. On a small cluster, it works fine. On a large cluster with CCS and a broad index pattern, this datafeed will scan every raw document on every run and never catch up with live data.</p><p>Instead of clicking <strong>Create</strong>, click <strong>Convert to advanced job</strong>. This keeps everything you just configured (the detector, the partition field, the bucket span) and drops you directly into the Advanced Wizard, where we can apply the five-step pattern.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt883a9c3cfbe84597/6a16fa0f839dfae25edcfce1/388d279820a639c08b119753f064b3a948ace8c6-1999x1591.png" alt="Kibana multi‑metric anomaly detection job summary showing a line chart split by geographic destination and a configuration panel with job settings, time range, and creation options." /><h3><strong>Analysis configuration</strong></h3><p>The conversion prefills the detector, partition field, and bucket span. The only change needed here is <strong>Step 1</strong> of the pattern: Open the <strong>Edit JSON</strong> view, and add <code>summary_count_field_name</code> to tell the ML engine that incoming data will be pre-summarized:</p>{
  "bucket_span": "1h",
  "summary_count_field_name": "doc_count", // Step 1
  "detectors": [
    {
      "function": "high_mean",
      "field_name": "bytes",
      "partition_field_name": "geo.dest"
    }
  ],
  "influencers": ["geo.dest"]
}<h3><strong>Datafeed configuration</strong></h3><p>Switch to the <strong>Datafeed</strong> tab. This is where Steps 2 through 5 of the pattern come together. Remove <code>scroll_size</code> if it’s present, and then enter the aggregations:</p>{
  "buckets": {
    "date_histogram": {               // Step 2: bucket wrapper, interval = bucket_span
      "field": "@timestamp",
      "fixed_interval": "1h"
    },
    "aggregations": {
      "@timestamp": {                 // Step 3: max timestamp anchor
        "max": { "field": "@timestamp" }
      },
      "geo.dest": {                   // Step 4: partition field, name must match exactly
        "terms": {
          "field": "geo.dest",
          "size": 1000
        },
        "aggregations": {
          "bytes": {                  // Step 5: metric field → avg aggregation
            "avg": { "field": "bytes" }
          }
        }
      }
    }
  }
}<p>A few notes on this config:</p><ul><li><p><strong>Step 2:</strong> The <code>date_histogram</code> uses <code>fixed_interval</code>: <code>"1h"</code>, matching <code>bucket_span</code> exactly. A mismatch produces incorrect bucket timing.</p></li><li><p><strong>Step 3:</strong> The <code>max</code> aggregation on <code>@timestamp</code> must be named <code>@timestamp</code> and placed inside the histogram’s <code>aggregations</code>; without it, the ML node can’t determine the precise end of each bucket.</p></li><li><p><strong>Step 4:</strong> The <code>terms</code> aggregation for the partition field must be named <strong>exactly</strong> after the partition field: <code>geo.dest</code>, not <code>geo.dest_grouping</code> or any alias. The ML engine uses the aggregation name, not the <code>field</code> parameter, to identify which partition value each bucket belongs to. A mismatch silently drops the partition field from results entirely.</p></li><li><p><strong>Step 5:</strong> The metric aggregation key <code>bytes</code> matches <code>field_name</code> in the detector exactly. Any mismatch here produces silently wrong anomaly scores.</p></li></ul><h3><strong>Validate with the preview</strong></h3><p>Before we create the job, let’s use the <strong>Preview</strong> tab. This runs the aggregation against real data and shows exactly what the ML node will receive, a very useful sanity check before committing.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1a6a2a786c4bb21a/6a16fa11a6c2b90ab5e794e7/638eeb2ae5b854195dec0e468887300b9afd2c58-1999x1254.png" alt="Three‑panel view showing an ML job configuration JSON, a matching datafeed JSON with aggregations, and a datafeed preview listing timestamped, bucketed results with fields like geo.dest, bytes, and doc_count." /><p>Three things to verify in the preview output: <code>doc_count</code> should be present on every bucket and greater than 1. The <code>bytes</code> values should look like average response sizes: numbers in the hundreds to hundreds of thousands for web traffic. And each row should correspond to a distinct (<code>timestamp</code>, <code>geo.dest</code>) pair. If anything looks off, fix it in the JSON editor and rerun the preview.</p><h2><strong>Adding influencer fields</strong></h2><p>In the example above, <code>geo.dest</code> is the partition field. The ML model learns a separate baseline for each destination country, and anomalies are reported per country. But you might also want <code>machine.os</code> to appear as an <strong>influencer</strong> in anomaly results: When the detector fires, you want to see “this looks anomalous for <code>geo.dest: CN</code> and <code>machine.os: win</code> is a contributing factor.” <a href="https://www.elastic.co/docs/explore-analyze/machine-learning/anomaly-detection/ml-ad-run-jobs#ml-ad-influencers">Influencers</a> don’t drive anomaly detection; they provide context for the anomalies that are found.</p><p>To support an influencer alongside a partition field, the analysis config gains an <code>influencers</code> array:</p>“Analysis_config”: {
  "bucket_span": "1h",
  "summary_count_field_name": "doc_count",
  "detectors": [
    {
      "function": "high_mean",
      "field_name": "bytes",
      "partition_field_name": "geo.dest"
    }
  ],
  "influencers": ["geo.dest", "machine.os"]
}<p>And now the datafeed needs to aggregate on both fields simultaneously. One <code>terms</code> nested inside another <code>terms</code> won’t work; a nested <code>terms</code> surfaces only the top-N values of the inner field per outer bucket, so you’d silently lose combinations. Instead, use a <a href="https://www.elastic.co/docs/reference/aggregations/search-aggregations-bucket-composite-aggregation">composite aggregation</a> with one <code>terms</code> source per field, next to the <code>date_histogram</code>:</p>"aggregations": {
    "buckets": {
      "composite": {
        "size": 1000,
        "sources": [
          { "timestamp": { "date_histogram": { "field": "timestamp", "fixed_interval": "1h" } } },
          { "geo.dest": { "terms": { "field": "geo.dest" } } },
          { "machine.os": { "terms": { "field": "machine.os.keyword" } } }
        ]
      },
      "aggregations": {
        "timestamp": { "max": { "field": "timestamp" } },
        "bytes": { "avg": { "field": "bytes" } }
      }
    }
  }<p><code>composite</code> generates one bucket per unique (<code>geo.dest</code>, <code>machine.os</code>) combination. The ML node sees every pair and can correctly attribute which operating system was contributing when a country’s response sizes spiked. Use the preview to confirm distinct pairs appear. If you only see a handful of rows where you’d expect many, the <code>size</code> parameter on the composite may need to be raised.</p><h2><strong>Categorization</strong></h2><p>Categorization works with aggregated datafeeds: <code>summary_count_field_name</code> and <code>categorization_field_name</code> can coexist in the same job. The five-step pattern applies directly. Step 2 uses the standard <code>date_histogram</code> topology. Step 4 has one adjustment: Instead of a partition field, we aggregate the text field itself using a <code>terms</code> aggregation on its <code>.keyword</code> subfield, named to match <code>categorization_field_name</code> exactly. Step 5 is skipped. The <code>count</code> detector works from <code>doc_count</code> alone.
<strong>Analysis config:</strong></p>{
  "bucket_span": "1h",
  "summary_count_field_name": "doc_count",
  "categorization_field_name": "message",
  "detectors": [
    {
      "function": "count",
      "by_field_name": "mlcategory"
    }
  ],
  "influencers": ["mlcategory"]
}<p><strong>Datafeed aggregations:</strong></p>{
  "buckets": {
    "date_histogram": {
      "field": "@timestamp",
      "fixed_interval": "1h"
    },
    "aggregations": {
      "@timestamp": {
        "max": { "field": "@timestamp" }
      },
      "message": {
        "terms": {
          "field": "message.keyword",
          "size": 1000
        }
      }
    }
  }
}<p>The datafeed sends one bucket per unique <code>message.keyword</code> value with a <code>doc_count</code> for each. The ML node receives those strings, runs categorization on them, assigning an <code>mlcategory</code> to each, and the <code>count</code> detector tracks how many documents fall into each category per bucket. The naming rule from Step 4 applies: The <code>terms</code> aggregation must be named <code>message</code>, matching <code>categorization_field_name</code> in the analysis config exactly.</p><p>One thing to watch: Keyword fields have a default <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/ignore-above"><code>ignore_above: 256</code></a> limit. Log messages longer than 256 characters won’t be indexed as <code>.keyword</code> and will be silently excluded from the aggregation. If your log messages are long, check the field mapping before using this approach. You may need to raise the limit in your index template.</p><h2><strong>The minimal pattern for </strong><strong><code>time_of_day</code></strong><strong> and </strong><strong><code>time_of_week</code></strong></h2><p><a href="https://www.elastic.co/docs/reference/machine-learning/ml-time-functions"><code>time_of_day</code></a><a href="https://www.elastic.co/docs/reference/machine-learning/ml-time-functions"> and </a><a href="https://www.elastic.co/docs/reference/machine-learning/ml-time-functions"><code>time_of_week</code></a> are the easiest functions to aggregate: They only need a timestamp and a document count. The C++ process extracts the time component from the bucket timestamp and builds a cyclical model of normal activity; <code>doc_count</code> tells it how many events fell in each bucket. No <code>terms</code> sources, no metric aggregation, no composite.
<strong>Analysis config:</strong></p>{
  "bucket_span": "15m",
  "summary_count_field_name": "doc_count",
  "detectors": [
    { "function": "time_of_day" }
  ]
}<p><strong>Datafeed aggregations:</strong></p>{
  "time": {
    "date_histogram": {
      "field": "@timestamp",
      "fixed_interval": "15m"
    },
    "aggregations": {
      "@timestamp": { "max": { "field": "@timestamp" } }
    }
  }
}<p>A plain <code>date_histogram</code> is enough; no composite needed. This makes <code>time_of_day</code> and <code>time_of_week</code> particularly CCS-friendly: one request per time chunk, minimal data over the wire. Use the same structure for <code>time_of_week</code>; only the function name changes.</p><p>If you want to add a <code>partition_field_name</code> (for example, to model time-of-day patterns per service), add a <code>terms</code> aggregation inside the histogram’s aggregations following the standard Step 4 pattern.</p><h2><strong>The composite pattern for </strong><strong><code>rare</code></strong><strong>, </strong><strong><code>freq_rare</code></strong><strong>, and </strong><strong><code>info_content</code></strong></h2><p><a href="https://www.elastic.co/docs/reference/machine-learning/ml-rare-functions"><code>rare</code></a><a href="https://www.elastic.co/docs/reference/machine-learning/ml-rare-functions">, </a><a href="https://www.elastic.co/docs/reference/machine-learning/ml-rare-functions"><code>freq_rare</code></a>, and <a href="https://www.elastic.co/docs/reference/machine-learning/ml-info-functions"><code>info_content</code></a> all need the composite extractor, the one that paginates through all unique value combinations rather than truncating to top-N. The five-step pattern applies here with a different topology in Step 2: <code>composite</code> goes at the top level (not <code>date_histogram</code>), with <code>date_histogram</code> as a source inside it. Step 3 places the <code>max</code> <code>@timestamp</code> aggregation as a sibling of the <code>composite</code>, and Step 5 is skipped since all three functions work from <code>doc_count</code> alone.</p><p>The datafeed structure is the same for all three functions: a composite at the top level, a <code>date_histogram</code> as one of its sources, and one <code>terms</code> source per analysis field. The only thing that varies is which fields you include as <code>terms</code> sources: <code>rare</code> needs one source for <code>by_field_name</code>; <code>freq_rare</code> needs sources for both <code>by_field_name</code> and <code>over_field_name</code>; <code>info_content</code> needs a source for <code>field_name</code> plus any <code>by_field_name</code> or <code>over_field_name</code> fields. None of the three require a metric aggregation.</p>{
  "buckets": {
    "composite": {
      "size": 10000,
      "sources": [
        { "@timestamp":   { "date_histogram": { "field": "@timestamp", "fixed_interval": "5m" } } },
        { "by_field":     { "terms": { "field": "by_field" } } },
        { "over_field":   { "terms": { "field": "over_field" } } }
      ]
    },
    "aggregations": {
      "@timestamp": { "max": { "field": "@timestamp" } }
    }
  }
}<p>A few notes:</p><ul><li><p>The composite aggregation must be the top-level aggregation, not nested inside a <code>date_histogram</code>. This is what routes the datafeed to the composite extractor.</p></li><li><p>The <code>date_histogram</code> is a source inside the composite, not the outer wrapper. Its <code>fixed_interval</code> must divide evenly into <code>bucket_span</code>.</p></li><li><p>The <code>max</code> aggregation on <code>@timestamp</code> sits as a sibling of the <code>composite</code> (inside <code>aggregations</code>), not nested inside it.</p></li><li><p><code>composite.size</code> controls the page size per round trip. Setting it high (10000) reduces round trips, which matters with CCS latency. With three sources and high-cardinality fields, the total combination count can be large; the extractor paginates automatically.</p></li></ul><h2><strong>Why aggregation-based datafeeds outperform scroll at scale</strong></h2><p>The gap is structural, not incidental. A scroll-based datafeed reads raw documents one page at a time: Every 1,000 documents is one request, and each waits for the previous one to complete before issuing the next. The number of requests is therefore proportional to the total document count in the time range being backfilled. At 836,000 events per hour over 13 months, that's roughly 7.9 billion events, or 7.9 million sequential round trips. Each round trip crosses the CCS boundary, waits for shard responses, and transfers matching documents in full. There’s no parallelism: The datafeed holds a <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-scroll">scroll context</a> open on the remote cluster and processes one page at a time.</p><p>An aggregation-based datafeed works differently. The data nodes summarize data locally, grouping by time bucket and categorical fields, and ship only the bucket results to the ML node. The number of requests is proportional to field cardinalities, not document count. In our example, two influencer fields with six unique combinations produce six result rows per time bucket; the datafeed pages through those in a handful of requests regardless of how many raw events fall in each bucket. Double the ingestion rate and the scroll request count doubles; the aggregation request count stays the same. This is why the gap widens at scale: The more data you have, the worse scroll looks by comparison, and the better aggregations look.</p><p>On live data, the picture is different because each real-time tick covers only one fresh bucket: Scroll issues however many pages fit in that bucket's worth of data, while aggregations issue one request. The 20× figure for live data reflects that ratio at 836,000 events per hour with a 15-minute bucket span. The practical threshold where aggregations stop being optional is when <code>(ingestion rate × bucket span) &gt; scroll_size</code>; once a single bucket contains more than <a href="https://www.elastic.co/docs/explore-analyze/machine-learning/anomaly-detection/anomaly-detection-scale#set-scroll-size">one scroll page</a> of documents, the datafeed can't keep pace with live data regardless of hardware. Below that threshold, scroll is fine and aggregations are a nice-to-have. Above it, aggregations are the only sustainable option.</p><p>Scroll-based datafeeds are the right default, and the wizards make the right call for most deployments. At scale (more shards, broader index patterns, CCS across tiers), switching to an aggregation-based datafeed is the natural next step: The data nodes summarize where the data lives, the ML node processes compact results, and the detections stay the same. The one cost to know up front is model state: Switching requires a new job, so the earlier you make the move, the less you give up.</p><p>If you hit a case not covered here, an aggregation type that doesn’t map cleanly or a composite that behaves unexpectedly, the <a href="https://discuss.elastic.co/">Elastic Discuss forums</a> are a good place to continue.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elastic-machine-leaning-jobs-aggregation-datafeeds</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elastic-machine-leaning-jobs-aggregation-datafeeds</guid>
    <category><![CDATA[ML Research]]></category>
    <dc:creator><![CDATA[Valeriy Khakhutskyy]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt584a9fa4d6ed3889/6a16fa13a6c2b97c15e794eb/023e3e6cb25891f789129d496c181113cc570f1f-1280x720.png" length="0" type="image/png"/>
    <pubDate>Wed, 15 Apr 2026 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>