<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Security Operations - Elastic Security Labs]]></title>
    <description><![CDATA[Trusted security news & research from the team at Elastic.]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Security Operations - Elastic Security Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte2c6b841aff36df4/6a88d9784acc96e3f324863d/security-labs-thumbnail.png</url>
      <link>https://www.elastic.co/security-labs/blog/category/security-operations</link>
    </image>
    <link>https://www.elastic.co/security-labs/blog/category/security-operations</link>
    <atom:link href="https://www.elastic.co/security-labs/rss/category/security-operations.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Mon, 14 Sep 2026 21:39:35 GMT</lastBuildDate>
  <item>
    <title><![CDATA[Data access: the hidden cost of security vendor lock-in]]></title>
    <description><![CDATA[Getting data into a security platform is always easy; getting it back out is where vendors add cost, extra tooling, and latency, and it is the part of the evaluation most teams overlook.]]></description>
    <content:encoded><![CDATA[<p>Your security data is the most important asset in your SOC. Not the dashboards, not the detections, not the AI features on the roadmap slide. The data. And most vendors make you pay, wait, or license your way to getting it back out. Every investigation your analysts run, every model you train, every agent you deploy is only as good as the telemetry underneath it. So here is the question I want you to ask every vendor in your stack: if I want my data, right now, what does that take?</p><p>Most vendors cannot answer it cleanly. And the answer matters more than almost anything else on the RFP.</p><h2><strong>What open actually means</strong></h2><p>The industry has gotten comfortable treating “open” as a checkbox. Support an open schema, publish an API, sponsor a standard, done. But a standard tells you how data is shaped. It tells you nothing about whether you can actually get it. Open data means three things, and you need all three:</p><p><strong>It is yours, at no additional cost.</strong> You generated this telemetry. You paid to collect it, and you paid to store it. If your vendor charges you a second time to access it, that is not a data feature. That is a toll booth on your own driveway.</p><p><strong>It is all of your data.</strong> Not the alerts. Not a normalized summary. Not the tables the vendor decided are ready. The full-fidelity record because you cannot predict today which field matters in next year’s investigation.</p><p><strong>It is real time.</strong> This one used to be a nice-to-have. Not anymore. <a href="https://www.crowdstrike.com/en-us/global-threat-report/">CrowdStrike’s own 2026 Global Threat Report </a>clocked the average eCrime breakout time at 29 minutes, the fastest at 27 seconds, and one intrusion where exfiltration started four minutes after initial access. Attackers are moving faster too, using AI to shorten the gap between breaking in and doing real damage. If your telemetry arrives in batches, minutes apart, the attack can be over before your data shows up. You would not accept a smoke detector that checks for fire twice an hour.</p><p>When a vendor fails any of these three, they are not protecting your data. They are building an artificial moat with it. Ingestion is always frictionless. Egress is licensed, delayed, or degraded. That asymmetry is not an accident of engineering. It is the business model, and the industry has a name for it: lock-in.</p><h2><strong>What the documentation actually says</strong></h2><p>Don’t take our word for it. Every claim in this table links to the vendor’s own documentation. Read it yourself, and hold us to the same bar.</p><p><strong>Vendor</strong></p><p><strong>How you get your data out</strong></p><p><strong>Cost to access your data</strong></p><p><strong>Freshness</strong></p><p><strong>Fidelity</strong></p><p><strong>Openness</strong></p><p><strong>CrowdStrike</strong></p><p><a href="https://developer.crowdstrike.com/accomplish/stream-and-analyze-data/">Falcon Data Replicator</a>: batched file export to object storage</p><p><a href="https://www.crowdstrike.com/en-us/resources/data-sheets/falcon-data-replicator/">Licensed add-on</a></p><p>Bulk batches, not a live stream; <a href="https://www.crowdstrike.com/en-us/blog/crowdstrike-falcon-and-humio-leverage-all-your-fdr-data-in-one-place/">deleted after 7 days</a> for CrowdStrike managed-buckets</p><p>Raw telemetry is only available via FDR; the <a href="https://developer.crowdstrike.com/accomplish/stream-and-analyze-data/">event stream API</a> sends detections, not telemetry</p><p><strong>Restricted</strong></p><p><strong>Palo Alto Networks</strong></p><p>XSIAM <a href="https://cortex-docs.paloaltonetworks.com/cortex-xsiam/configure-cortex-xsiam/data-management/manage-event-forwarding">Event Forwarding</a>: batch files to a Palo Alto-managed bucket</p><p><a href="https://cortex-docs.paloaltonetworks.com/cortex-xsiam/learn-about-cortex-xsiam/cortex-xsiam-product-licenses">Two paid Event Forwarding add-ons</a>: GB and Endpoint</p><p><a href="https://cortex-docs.paloaltonetworks.com/cortex-xsiam/configure-cortex-xsiam/data-management/manage-event-forwarding">Batch; up to 2 hours to appear; kept 14 days</a></p><p>Endpoint and log data <a href="https://cortex-docs.paloaltonetworks.com/cortex-xsiam/learn-about-cortex-xsiam/cortex-xsiam-product-licenses">split across the two add-ons</a></p><p><strong>Restricted</strong></p><p><strong>Microsoft</strong></p><p>Defender <a href="https://learn.microsoft.com/en-us/defender-xdr/streaming-api">streaming API</a>; Sentinel <a href="https://learn.microsoft.com/en-us/azure/azure-monitor/logs/logs-data-export">data export</a></p><p>Streaming: pay Azure infra; <a href="https://learn.microsoft.com/en-us/azure/azure-monitor/logs/logs-data-export">Sentinel export billed per GB</a></p><p><a href="https://learn.microsoft.com/en-us/defender-xdr/api-overview">Real-time stream</a> (via Azure Event Hubs)</p><p><a href="https://learn.microsoft.com/en-us/defender-xdr/supported-event-types">Generally-available fields only</a>; Azure destinations only</p><p><strong>Partially open</strong></p><p><strong>Google</strong></p><p><a href="https://docs.cloud.google.com/chronicle/docs/reference/data-export-api-enhanced">Bulk export</a> to your storage, or a <a href="https://docs.cloud.google.com/chronicle/docs/reports/bigquery-export">continuous BigQuery feed</a></p><p>Export capped; <a href="https://docs.cloud.google.com/chronicle/docs/reports/bigquery-export">BigQuery billed by query</a></p><p><a href="https://docs.cloud.google.com/chronicle/docs/reports/bigquery-export">Live feed 5-10 min behind</a> (top tier only); raw export is point-in-time</p><p>Live feed is normalized only; no continuous raw path</p><p><strong>Limited</strong></p><p><strong>Splunk </strong></p><p><a href="https://help.splunk.com/en/splunk-cloud-platform/search/search-manual/9.3.2411/export-search-results/export-data-using-the-splunk-rest-api">Search/export REST API</a></p><p>Included</p><p>On-demand query via <a href="https://help.splunk.com/en/splunk-cloud-platform/search/search-manual/9.3.2411/export-search-results/export-data-using-the-splunk-rest-api">export data API</a>; data available near-real-time</p><p>Full events returned as structured JSON</p><p><strong>Open</strong></p><p><strong>Elastic</strong></p><p>REST APIs including <a href="https://www.elastic.co/docs/solutions/search/the-search-api">Search</a>, <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-open-point-in-time">Point in time</a> and <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-rest">ES|QL</a></p><p>No export license or per-query export fee; <a href="https://www.elastic.co/docs/deploy-manage/cloud-organization/billing/cloud-hosted-deployment-billing-dimensions">Elastic Cloud applies ordinary cloud data-transfer metering</a>, not an egress toll</p><p>On-demand query via API; data available to query <a href="https://www.elastic.co/docs/manage-data/data-store/near-real-time-search">as soon as it is indexed</a></p><p>Full raw and parsed events, returned as structured JSON</p><p><strong>Open</strong></p><p><em>This table reflects each vendor's public documentation as of September 2026.</em></p><p></p><p>A note on the ratings, because we want them to be defensible, not convenient. Open means all three tests pass: no added cost, full fidelity, and no batch window between the moment data arrives and the moment you can use it. Data you can query the moment it lands clears that bar; a scheduled batch measured in minutes or hours does not. Partially open means the vendor genuinely tries but with real constraints; credit to Microsoft for a true streaming path, even if it covers only generally-available fields and only lands in Azure. Limited means the data comes back, but late or in a form only the vendor can read. Restricted means access to your own telemetry is a paid product, a delayed batch, or both.</p><p>And since Elastic is on this list too, as an endpoint agent and a SIEM, hold us to the same bar. Our rating is about live access: APIs that return JSON, with no license between you and your own telemetry. On Elastic Cloud you pay ordinary data-transfer rates like any cloud service. What you never pay is a license to reach data that was already yours.</p><p>If any vendor believes we have mischaracterized their documentation, I genuinely want to hear it, and we will correct it.</p><h2><strong>So what does this actually cost you?</strong></h2><p>Here is what that comparison means when it really counts, in the middle of an incident. Detection you cannot act on in time is not detection. When your telemetry arrives in scheduled batches, a window opens between the alert and the evidence. Sometimes minutes, sometimes longer, and in that window the attack is live while your data is not yet in front of you. You paid to collect that telemetry. You paid to store it. And at the one moment it matters, it is still in transit. Or worse, it is not there at all because exporting it requires a separate license. That’s the difference between stopping an intrusion and reading about it afterward.</p><p>You do not create that gap during an incident. You inherit it at purchase. You already run a strong endpoint agent. It works, and you trust it. Now that same vendor offers to be your SIEM as well. One console, one relationship, one invoice. There is nothing wrong with one vendor doing both. We do both at Elastic. The question is what it costs you to change your mind. So look closely at what it takes to move that endpoint telemetry somewhere other than their own platform. Every vendor makes it easy to get data in. Getting it back out is something you buy, then wait for.</p><p>It is rarely just one handoff. Most security environments are heterogeneous by design, with each layer chosen because it is good at its own job. Modern attacks move across all of them, from a stolen identity to a cloud workload to an endpoint, and you only see the full chain if the data from each can meet in one place, quickly and without a toll. When a vendor makes its telemetry expensive or slow to share, it is not just charging you. It’s fragmenting the picture, and a fragmented picture is exactly where intrusions hide. The all-in-one pitch offers to solve that. But the fragmentation was manufactured: vendors made their data hard to move, then sold you the one console where it comes back together. So before convenience makes the decision for you, ask three questions, and make the vendor answer them in writing:</p><ul><li><p>Can I get all of my telemetry out, continuously, without buying a separate license to do it?</p></li><li><p>How long, exactly, from the moment an event happens to the moment it is usable in a system I chose?</p></li><li><p>On the day I add a different analytics engine, a data lake, or a new AI model, what breaks?</p></li></ul><p>If the honest answers are “extra cost,” “in batches,” and “quite a lot,” then you are not buying a SIEM. You are renting access to your own data. The point is not that these layers must stay separate. Plenty of teams consolidate for good reasons, and we sell both layers ourselves. The point is that the choice should stay yours: you can add the analytics platform you want, or leave the one you have, without paying a toll or waiting on a batch to get your own data. The only reason to accept less is that a vendor made leaving hard enough that staying felt like a decision. It was not a decision. It was the absence of one.</p><h2><strong>The strategic question</strong></h2><p>Your SIEM is not a tool you bought. It is where isolated alerts become an attack story, and where you go to find out what actually happened. The vendor holding it is a strategic dependency, whether you planned it that way or not.</p><p>So evaluate them like one. Put data portability on the RFP, not in the demo, and score it on two things: what it costs to move your telemetry somewhere else, and how much delay it adds.</p><p>Because you cannot build a real-time defense on a delayed copy of your own telemetry, and you should not have to buy your data back to try.</p><p>It is your data. Any vendor who makes that complicated has told you what kind of partner they intend to be.</p>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/siem-data-export-comparison</link>
    <guid isPermaLink="false">siem-data-export-comparison</guid>
    <category><![CDATA[Security Operations]]></category>
    <category><![CDATA[SOC]]></category>
    <dc:creator><![CDATA[Mike Nichols,Jamie Hynds]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt194a2b93239a98c3/6a9a90363481c2bd668af9b3/2189.png" length="0" type="image/png"/>
    <pubDate>Fri, 04 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[How a team of entity maintainers monitors, connects and scores entities in Elastic Security]]></title>
    <description><![CDATA[Inside Elastic Security, background jobs called maintainers each own one piece of every user, host and service record, from building entities out of raw logs to resolving identities and scoring risk.]]></description>
    <content:encoded><![CDATA[<p>Open the entity analytics (EA) graph in Elastic Security and you'll see a user wired to the hosts they log in to and the devices they own, along with scattered accounts that turn out to be the same user. While interesting on its own, it provides a critical piece of context during a threat hunting or incident investigation. This post gives an overview of EA fundamentals and opens the hood to see how edges are drawn and accounts are resolved. Why is that important? Well, everything downstream, including baselines, risk, and AI reasoning, is only as good as the entity records underneath it.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta456ecc712e7bbc6/6a8e952a1106e8de8d9c4b0d/image1.png" alt="Figure 1: Entity graph surfacing relationships." title="Figure 1: Entity graph surfacing relationships." /></p>
<p>Before we go deeper, let's set some context. Many SIEMs treat entities as flat records, a snapshot of what's true right now, rebuilt on demand from raw logs. That works until you need to know how the environment got here, what changed, what was resolved, how a risk score compounded over time. Elastic's Entity Store is architected differently. Every entity is a living record, continuously enriched by background jobs called <em>maintainers</em>, each responsible for one facet of the entity. Maintainers establish relationships, perform identity resolution, and calculate risk scoring. These are all composed onto the record over time, all inspectable, and all correctable. A companion post on <a href="https://www.elastic.co/security-labs/ueba-entity-record-quality-analytics">entity record quality</a> explains why this matters.</p>
<p>Going one level deeper, this piece explores <em>how</em> records are built and connected in the first place, along with <em>how</em> their risk scores are updated, relationships are tracked, and identities are resolved. The short answer is a single abstraction. Entity store v2, introduced in Elastic v9.4, is built on maintainers. Each maintainer runs on its own clock, continuously updating the record you eventually query. The entity store is the result of a set of maintainers enriching each entity based on the raw or derived signal around it, not an append-only index you simply write entities into.</p>
<p>Let's build that picture from the engine up.</p>
<h2 id="whattheentitystoredoesandwhyitrunsonesql">What the entity store does and why it runs on ES|QL</h2>
<p>The Entity Store is the layer that turns your telemetry (for example endpoint, identity provider, and cloud logs), identity and asset inventories, and threat detections into one queryable profile per user, host, and service. This means analysts can pivot on information and insights gained from an enriched entity record, instead of reconstructing it from raw logs every time.</p>
<p>Entity store Version 1 did that with <a href="https://www.elastic.co/docs/explore-analyze/transforms/transform-overview">Elasticsearch transforms</a>.</p>
<p>Version 2 rebuilds the engine on <a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language (ES|QL)</a>. Each entity extraction cycle runs a query that filters the relevant logs and aggregates them. In the next crucial step, the cycle performs a <code>LOOKUP JOIN</code> back against the existing entity index to carry forward previously observed entity attributes and behaviors.</p>
<pre><code>FROM logs-*
  | WHERE ...                                          // logs for this entity type + window
  | STATS ... BY entity.id                             // collapse many events into one record
  | LOOKUP JOIN .entities.v2.latest.security_default-00001
      ON entity.id                                     // field retention against the store itself
  | EVAL ...                                            // keep-latest / keep-first retention
</code></pre>
<p>ES|QL provides query flexibility and coverage for the fields and field types that the schema needs. It also provides a pipeline that you can reason about and extend, with retention and merge logic expressed directly, instead of through the limited flexibility and operational choreography needed with transforms.</p>
<h3 id="howanentityrecordisbuiltfromrawsignaltoresolvedidentity">How an entity record is built, from raw signal to resolved identity</h3>
<p>Here's the part that's usually glossed over. Let's take a look at the entity extraction and building process by following one host from its raw event to its finished record.</p>
<ul>
<li><p><strong>Step 1: The raw signal ingested to Elasticsearch.</strong> Your endpoint detection and response (EDR) agent and identity provider write events into data streams that are matched by a <code>logs-*</code> index pattern, as do your cloud integrations. There isn’t anything entity-shaped yet, just events with fields like <code>user.name</code>, <code>host.name</code>, and <code>event.category</code>.</p></li>
<li><p><strong>Step 2: Extraction collapses events into a record.</strong> During its run cycle, the ES|QL query above filters the logs relevant to an entity definition, and then <code>STATS ... BY entity.id</code> collapses potentially thousands of events into one entity store row/record per entity, keeping the first-seen and latest entity lifecycle values that matter. The <code>LOOKUP JOIN</code> merges that fresh row with the entity's existing record, so nothing observed outside the window is lost. The output is a single, denormalized document in the latest index: the entity's base identity and lifecycle, in addition to its attributes.</p></li>
<li><p><strong>Step 3: Identity is keyed deliberately.</strong> Which entity a row belongs to is the highest-stakes decision in the whole system. Key it wrong and you can blend several distinct users into one entity, or split one user across many. So the store doesn't guess. It derives a deterministic identifier, the entity unique ID (EUID), from the fields that actually identify the entity.</p></li>
<li><p><strong>Step 4. Maintainers compose the rest.</strong> The entity record is there, but it doesn’t yet know who it talks to (relationships) or what behaviors it exhibits. It also doesn’t yet know which various account entities are used by the same user or how risky the entity is. Those facets are layered onto the entity record by maintainers, and that's the heart of Entity Analytics.</p></li>
</ul>
<h3 id="thekeytobuildingentityrelationshipsthedeterministicentityid">The key to building entity relationships - the deterministic entity ID</h3>
<p>The Entity Store currently supports three primary, schema-backed entity types: <strong>users, hosts, and services</strong>. Alongside these sits a fourth, polymorphic category known as <em>generic</em> entities.<br />
The <em>generic</em> entity type isn’t constrained to a fixed schema, but rather provides an extensibility layer for the entity store platform. For example, by leveraging cloud and orchestrator fields, it represents resources like <strong>EC2 instances, S3 buckets, and Kubernetes clusters</strong> as first-class entities today. This flexibility allows the entity store to encompass a broad spectrum of assets without the architectural overhead of defining a new entity type for every resource class.</p>
<p>Everything the entity store does, including merging today's events onto the current entity record and linking two accounts into one resolved user, hinges on one field: <code>entity.id</code>, the EUID.</p>
<p>The EUID is a deterministic, human-readable string (a type prefix and then the fields that actually identify the entity, joined with <code>@</code>), rather than a random universally unique identifier (UUID) or a hash. Note the addition of a user entity <em>namespace</em> at the end of the user EUID. We’ll explain the role that plays below.</p>
<pre><code>user:jane.doe@example.com@okta
host:9f86d081-1e0c-4b3f-8a2d-2c1e7bed425e
service:api-gateway
</code></pre>
<p><strong>How is the EUID derived?</strong> For each entity, the store walks an ordered list of candidate identity fields and takes the first <em>complete</em> one. If a field a candidate needs is missing, that candidate is skipped and the next is tried, so a partial observation never produces a malformed ID.</p>
<p><strong>Why do we need a namespace for user entities?</strong> While host and service entities rely on sufficiently unique keys like <code>host.id</code> or <code>service.name</code>, a user's identity, for example an email address, can be observed in many different log sources, but is authoritative only within its issuing domain. The namespace provides the necessary disambiguation, resolving the specific collision challenges unique to user identity records.</p>
<p>| Entity type              | ID fields (priority order)                                      | Example EUID                                           | Namespace                                                      |
| :----------------------- | :-------------------------------------------------------------- | :----------------------------------------------------- | :------------------------------------------------------------- |
| Host                     | host.id, then host.name, then host.hostname                     | <code>host:9f86d081-1e0c-4b3f-8a2d-2c1e7bed425e</code>            | (none)                                                         |
| Service                  | service.name                                                    | <code>service:api-gateway</code>                                  | (none)                                                         |
| User (identity provider) | user.email, then user.id, then user.name@domain, then user.name | <code>user:jane.doe@elastic.com@okta</code>                       | Provider name: okta, entra_id, microsoft_365, active_directory |
| User (local/endpoint)    | user.name scoped to host.id                                     | <code>user:jdoe@9f86d081-1e0c-4b3f-8a2d-2c1e7bed425e@local</code> | local                                                          |</p>
<ul>
<li><strong>Host:</strong> The first present of <code>host.id</code>, <code>host.name</code>, <code>host.hostname</code>, giving <code>host:&lt;value&gt;</code>.</li>
<li><strong>Service:</strong> <code>service.name</code>, giving <code>service:&lt;name&gt;</code>.</li>
<li><strong>User:</strong> This is the interesting one, because a user's identity depends on <em>where their activity was observed.</em></li>
</ul>
<p><strong>Users: Identity provider versus the local host.</strong> A user's EUID always ends in a <em>namespace</em>, the last <code>@</code>-delimited segment, and that segment is what stops two accounts that merely <em>look</em> alike from colliding.</p>
<p>When a user comes from an identity provider, the namespace is that provider (<code>okta</code>, <code>entra_id</code>, <code>microsoft_365</code>, <code>active_directory</code>) and the store keys on the most authoritative identifier it has: email, then user id, then <code>name@domain</code>, and then name.</p>
<pre><code>user:jane.doe@elastic.com@okta
</code></pre>
<p>The same user's Entra ID account becomes <code>…@entra_id</code>, a <em>different</em> EUID, on purpose. They’re two authoritative accounts until resolution ties them together. We are going to see how resolution achieves this later in the blog.</p>
<p>When a user is seen only through endpoint or host telemetry, with no authoritative directory account behind them, the store scopes them to the machine and suffix the ID as <code>local</code>:</p>
<pre><code>user:jdoe@9f86d081-1e0c-4b3f-8a2d-2c1e7bed425e@local
</code></pre>
<p>The middle segment is the host's <em>durable</em> identifier (<code>host.id</code>) and not its renamable hostname, so <code>jdoe</code> on a laptop stays distinct from <code>jdoe</code> on a shared bastion host, and the identity survives a machine being renamed or reimaged. The <code>local</code> namespace is an explicit signal: <em>This is activity on this box, not a verified global identity.</em> (Note the host in the example above and this local user share the same <code>host.id</code>. That's how the Entity Store knows they belong together.)</p>
<h3 id="whatmakesanentitysource_authoritative_andhowtoconstructone">What makes an entity source <em>authoritative</em> and how to construct one</h3>
<p>The EUID logic leans on a word that deserves a precise definition: <em>authoritative</em>. When an identity-provider-backed user earns a real namespace (<code>okta</code> or <code>entra_id</code>, among others) and high-confidence treatment, it's because the incoming events cleared a specific bar. Here's the bar and how to clear it with your own integrations.</p>
<p><strong>What qualifies as authoritative identity data.</strong> The store treats an event as an authoritative identity signal when either of these is true:</p>
<p>| Classification path                     | Required ECS fields                                       | Confidence | Resulting namespace                  |
| :-------------------------------------- | :-------------------------------------------------------- | :--------- | :----------------------------------- |
| Asset inventory                         | event.kind: asset                                         | High       | Provider name (okta, entra_id, etc.) |
| Endpoint-observed (neither bar cleared) | user.name + host.id present, but no authoritative signal | Medium     | local (scoped to host)               |
| Unclassifiable                          | None of the above                                         | N/A        | No entity created                    |</p>
<ul>
<li>The event is an <em>asset / inventory document</em> with <code>event.kind: asset</code>. This is how Elastic's identity integrations (Okta, Entra ID, Active Directory, and the cloud asset sources) publish their user and account inventories.</li>
</ul>
<p>Clear either bar and the user becomes a first-class, identity-provider-backed entity at <em>high</em> confidence, and the preferred canonical record when identities are later resolved. Miss both, but carry a <code>user.name</code> and a <code>host.id</code>, and the user is instead scoped to that host in the <code>local</code> namespace at <em>medium</em> confidence. If both of these conditions are false and no host exists to associate the event with, no entity is created at all. That last case is deliberate restraint, not a gap; the store would rather create nothing than manufacture a noisy identity from an ambiguous event.</p>
<p>Two filters apply before any of this runs: The event's <code>event.outcome</code> must not be <code>failure</code> (a failed login isn’t evidence that an account exists), and it must carry at least one of <code>user.email</code>, <code>user.id</code>, or <code>user.name</code>.</p>
<p><strong>Which namespace you get.</strong> Once an event qualifies, the namespace is derived from its source, the first non-empty of <code>event.module</code> or the leading segment of <code>data_stream.dataset</code>.</p>
<p>An authoritative event from an <em>unrecognized</em> source still creates a real entity. It just lands in the <code>unknown</code> namespace and you lose provider-level disambiguation (identical usernames from two namespaces can collide) and clean grouping. So, for a custom integration, the goal isn't only to <em>qualify</em> as authoritative, it's to be <em>recognized</em>.</p>
<p><strong>How to accommodate a custom integration or pipeline.</strong> If you ingest identity data through a custom Fleet integration, a Logstash pipeline, or an Elasticsearch ingest pipeline, set these Elastic Common Schema (ECS) fields so the store classifies your entities the way you intend:</p>
<ol>
<li><strong>Mark the shape.</strong> For an account inventory, such as a local directory service, or configuration management database (CMDB), set <code>event.kind: asset</code>.</li>
<li><strong>Provide an identity.</strong> Populate at least one of <code>user.email</code> (this is best, since it's the top priority), <code>user.id</code>, or <code>user.name</code>, plus <code>user.domain</code> where you have it.</li>
<li><strong>Name your source.</strong> Set <code>event.module</code> (or the leading segment of <code>data_stream.dataset</code>) to a value the store maps. If you're feeding one of the known providers, reuse its naming so you inherit the right namespace. If your source is genuinely new, expect <code>unknown</code> until a mapping is added.</li>
<li><strong>Don't let real identities look local.</strong> The <code>local</code> classification triggers when an identity event carries both <code>user.name</code> and <code>host.id</code> but <em>doesn't</em> clear the authoritative bar. If you're publishing directory data, don't attach a <code>host.id</code> to it. That field is the signal that says "endpoint-observed, scope it to this box." Reserve it for genuinely host-local activity.</li>
<li><strong>Skip the noise.</strong> Don't emit <code>failure</code> outcomes as identity evidence. Note, too, that the store already excludes common shared and service account names (<code>root</code>, <code>jenkins</code>, <code>deploy</code>, <code>postgres</code>, <code>admin</code>, and similar) from the <code>local</code> namespace, so they never become per-host user entities.</li>
</ol>
<p>Get these right, and your custom source behaves exactly like a built-in one: Authoritative users resolve and score at full fidelity, and endpoint-observed users stay correctly host-scoped. Plus, nothing downstream has to special-case where the data came from.</p>
<h2 id="howmaintainersaddentityresolutionrelationshipsandriskscoringtoentities">How maintainers add entity resolution, relationships, and risk scoring to entities</h2>
<p>As discussed at the start of this post, the entity store is updated by dedicated background tasks called maintainers. Each of these has a specific job: building relationships between entities and resolving identities, along with updating risk scores.</p>
<p>The framework gives every maintainer its lifecycle, scheduling, and health reporting for free, making the entity store <em>extensible by design.</em> Any new entity enrichment capability ships as a new maintainer on the same rails, without re-architecting the entity store or breaking backward compatibility. Several maintainers already run in the current version, and recent additions, such as entity relationships derived from observed entity behaviors, have been added exactly this way. Even entity risk scoring, one of the most impactful risk-centric capabilities in the product, is implemented as just another maintainer composing one more facet of the entity record.</p>
<h3 id="maintainersdiscoverandstoreentityrelationships">Maintainers discover and store entity relationships</h3>
<p>Let’s take a look at one of the relationship maintainers. Once a day, the <code>accesses_frequently</code> / <code>accesses_infrequently</code> maintainer runs an ES|QL query over relevant telemetry and, for each actor→target pair, counts successful accesses over a 30-day window. If the count is:</p>
<ul>
<li>greater than or equal to an Elastic-defined threshold, the relationship becomes <code>accesses_frequently</code>.</li>
<li>less than the threshold, the relationship becomes <code>accesses_infrequently</code>.</li>
</ul>
<p>Today, it reads from Elastic Defend (endpoint logins), AWS CloudTrail (<code>StartSession</code> / <code>SendSSHPublicKey</code>), system auth (SSH logins), and system security (Windows events <code>4624</code>/<code>4648</code>). A sibling, <code>communicates_with</code>, builds communication links from Elastic Defend, system auth, system security, Jamf Pro, and AWS CloudTrail.</p>
<p>The relationships are written straight onto the entity record as arrays of target IDs:</p>
<pre><code>// a user entity in .entities.v2.latest.security_default-00001
"entity": {
  "id": "user-abc…",
  "relationships": {
    "accesses_frequently": { "ids": ["host-def…", "host-ghi…"] },
    "communicates_with":   { "ids": ["service-jkl…"] }
  }
}
</code></pre>
<h3 id="amaintainerperformsentityresolution">A maintainer performs entity resolution</h3>
<p>The resolution maintainer links fragmented accounts, the Okta <code>jdoe</code>, the Entra ID <code>jdoe</code>, and the on-prem Active Directory <code>jdoe</code>, into one resolved user. The risk maintainer then scores the resolved user, so the risky service token and the benign laptop login that both belong to John Doe are scored as John Doe: one record, not three that an analyst has to reconcile in their head.</p>
<p>Resolution happens two ways. Most of it is automatic: a background task runs every five minutes and links user entities that share the same <code>user.email</code>, so accounts converge on their own as the data arrives. When you need to step in, you can link or unlink manually from the entity flyout in the UI, or call the API directly.</p>
<pre><code>POST /api/security/entity_store/resolution/link
{
  "target_id": "user:jane.doe@elastic.com@okta",
  "entity_ids": ["user:jane.doe@elastic.com@entra_id"]
}
</code></pre>
<pre><code>POST /api/security/entity_store/resolution/unlink
{
  "entity_ids": ["user:jane.doe@elastic.com@entra_id"]
}
</code></pre>
<p>The near-term plan is to expand the out-of-the-box matching logic so more identities link automatically without anyone manually creating the links.</p>
<h3 id="amaintainercalculatesandupdatesentityriskscores">A maintainer calculates and updates entity risk scores</h3>
<p>The risk score maintainer runs hourly by default. On each run, it queries the detection alerts over a configurable rolling time window (past 30 days by default) and for each entity those alerts touch, aggregates the alert risk scores into a single normalized entity risk score and risk level.</p>
<p>Entity risk scoring also extends across resolved identities. First the risk scoring maintainer calculates an entity risk score for each user entity (one per account) based on any detection alerts related to that account. Next it calculates a distinct entity risk score for the <em>resolved</em> user across all of their linked accounts. For example, a malicious service token alert on one account, and an excessive laptop login rate alert from a different account, where both accounts belong to John Doe, roll up into a single user entity risk score for the resolved John Doe. That holistic view is possible because the resolution maintainer ran first, then the risk scoring maintainer reads the graph (set of linked accounts) that the resolution maintainer produced, and is able to assign a risk score to the user, not only the constituent accounts If an analyst or AI agent decides to investigate the provenance of a resolved user risk score, the answer is easily obtained.</p>
<h2 id="everyentitystoredecisionisinspectableandcorrectable">Every entity store decision is inspectable and correctable</h2>
<p>The store's engine and its relationship maintainers are built on ES|QL, the same query language that Elastic users already employ, so what a maintainer reads,computes, and writes is inspectable, rather than proprietary magic. Every entity lives as a document in an Elasticsearch index, so its full state is there to read, including the relationships that were drawn and the accounts that were resolved into one user, plus the signals behind a risk score.</p>
<p>Entity data is wrong sometimes; that's the reality of identity data. Because the state is represented as Elasticsearch documents, you can correct them through APIs. When resolution gets it wrong, two people merged or one user split across two records, you can link or unlink identities right through the UI, and the risk score maintainer rescores the corrected resolved user on its next run. When you need to pull an entity into scope or set its criticality, <a href="https://www.elastic.co/docs/solutions/security/advanced-entity-analytics/watchlists">watchlists</a> let you say so directly. <a href="https://www.elastic.co/security-labs/entity-analytics-agent-builder">With Elastic Agent Builder</a>, an AI agent can walk you through an entity's data and make those corrections for you, without leaving the Elastic UI.</p>
<h2 id="whatentitymaintainersmeanforsecurityanalysts">What entity maintainers mean for security analysts</h2>
<p>With entity resolution, relationship mapping, and risk scoring running as automated maintainers, human analysts and AI agents gain efficient access to the entity context they need to perform investigations. The entity store is extensible, so entities continue gaining relationships and enrichments as new maintainers ship or new data sources are ingested, with no migration or re-architecture. Nothing is hidden, so every risk score traces back to the alerts and other factors (for example, asset criticality, watchlist membership) that produced it, and a finding becomes something you can explain using the Agent Builder chat with the entity analytics tools and skills rather than a black box you have to trust.</p>
<p>In addition, wrong data doesn't have to stay wrong; a bad resolution link takes one API call or Agent Builder action to fix. What you're left with is a single queryable record per entity, host, and service that can evolve over time.</p>
<h2 id="whatscomingnextintheentityanalyticsroadmap">What’s coming next in the Entity Analytics Roadmap</h2>
<p>The science of entity analytics is ongoing. One area we’re digging into is integrating non-human identities (NHI) as entities. AI agents, service accounts, and agentic workloads are among the fastest-growing concerns across today’s attack surface. While this research continues, the core objectives involve finding an optimal way to link an AI agent's sessions to existing entities by tracking its utilized Service accounts or connected devices and services, as well as assessing the risk of AI agent behavior</p>
<p>Also under the hood, we're reworking log extraction to split the work by data source confidence, so authoritative identities and lower-confidence enrichment are handled separately.<br />
Additionally, we’re exploring enhancing the <a href="https://www.elastic.co/security-labs/proactive-threat-hunting-ai-generated-leads">agentic UEBA capabilities</a> to reason about an entity, connect the dots, and suggest relevant leads for the analyst to explore.</p>
<p>Entity risk scoring is getting continued attention too. Currently, a detection alert carries one static alert risk score that affects the entity risk score of every entity it involves, so the actor and the target can come out looking equally risky. The direction is towards a dynamic, entity-centric risk score that gives each entity its own contribution based on the role it played in the observed interaction, and its own context.</p>
<p>For practitioners with their hands on the Elastic Security UI, the entity analytics overview experience is due for a refresh, with risky entities shown more simply and clearer actions to investigate them.</p>
<p>Entity analytics is available in Elastic Security. Learn more about <a href="https://www.elastic.co/docs/solutions/security/advanced-entity-analytics">entity analytics</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/entity-resolution-identity-scoring-elastic-security</link>
    <guid isPermaLink="false">entity-resolution-identity-scoring-elastic-security</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Uri Weisman,Mike Paquette]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta456ecc712e7bbc6/6a8e952a1106e8de8d9c4b0d/image1.png" length="0" type="image/png"/>
    <pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[The security signal log tailing can't see: tracking npm cooldown removals with Elastic Agent]]></title>
    <description><![CDATA[A 40-line CEL integration snapshots .npmrc files every 6 hours to catch cooldown removals. This post walks through the three ways we broke filestream before landing on snapshot semantics.]]></description>
    <content:encoded><![CDATA[<p>npm's <code>min-release-age</code> setting tells npm to ignore any package version published less than a set number of days ago, keeping freshly compromised releases out of <code>npm install</code> during the window when they do the most damage. Getting the setting onto developer workstations is straightforward. Knowing when someone quietly deletes it is a different problem entirely, and log-tailing inputs are no help because they only fire when lines are appended to a file. We built a ~40-line Common Expression Language (CEL) integration in <a href="https://www.elastic.co/elastic-agent">Elastic Agent</a> that snapshots every <code>.npmrc</code> on a 6-hour heartbeat. When <code>min-release-age</code> disappears from the next snapshot, the pipeline marks it <code>cooldown.absent = true</code>. This post walks through that pipeline, the filestream approach we tried first, and what we learned from three iterations before landing on the final design.</p>
<p>The full path from config file to dashboard looks like this:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blteb5ec3f717e84e68/6a7d839133fa8a2eca1ff9b8/vertical_flowchart.png" alt="npm's cooldown pipeline" title="npm's cooldown pipeline" /></p>
<h2 id="whydeveloperworkstationsarethenpmsupplychaingap">Why developer workstations are the npm supply-chain  gap</h2>
<p>CI/CD pipelines have their own supply-chain guardrails that can be managed at the enterprise level. Workstations are where the gap usually sits. A developer running <code>npm install</code> on a freshly compromised package is a different surface from the build pipeline, and the package manager itself is the right place to apply a delay before a fresh release becomes installable. npm's <a href="https://docs.npmjs.com/cli/v11/using-npm/config"><code>min-release-age</code></a> setting, available since npm 11.10, does exactly that: the value is a number of days, and npm excludes any version published more recently from resolution. In MITRE ATT\&amp;CK terms, the technique it blunts is <a href="https://attack.mitre.org/techniques/T1195/001/">Compromise Software Dependencies and Development Tools (T1195.001)</a>.</p>
<p>Most other package managers have an equivalent flag; <a href="https://cooldowns.dev">cooldowns.dev</a> tracks which ones, what each flag is called, and recommended values. This post focuses on npm, and recent attacks have happened, but the technique applies to any of them.</p>
<p>We enforce the setting with Jamf: a script writes <code>min-release-age=7</code> to each user's <code>.npmrc</code> and the machine-global npmrc, and re-applies once a day. A removal therefore self-heals within 24 hours, and the telemetry is what makes the removal visible at all: the 6-hour heartbeat catches the gap window before Jamf closes it, and shows us which hosts keep removing the setting.</p>
<h2 id="whatdoesannpmcooldownconfigfilelooklike">What does an npm cooldown config file look like?</h2>
<p>The target is small: npm writes the setting to a user-level <code>.npmrc</code>, typically 10 to 50 bytes.</p>
<pre><code>~/.npmrc    →    min-release-age=7        # days
</code></pre>
<p>Machine-global settings live in a handful of well-known paths (<code>/opt/homebrew/etc/npmrc</code> and <code>/usr/local/etc/npmrc</code> on macOS, <code>/etc/npmrc</code> and <code>/usr/lib/node_modules/npm/.npmrc</code> on Linux), so a complete inventory has to include those paths too. While we haven’t deployed a Windows version yet, we did include all the details to make this work in the Linux and Windows variants section. These can be found in the section Linux and Windows variants.</p>
<p>Two ingredients made this look easy. First, as InfoSec is Customer Zero at Elastic, the Elastic Agent is already rolled out to every endpoint and runs with the privileges to read these files. Second, the schema is essentially flat: one key, one value, one config file. So in theory, you configure a file input, tail the cooldown lines, and ship them to <a href="https://www.elastic.co/elasticsearch">Elasticsearch</a>.</p>
<p>The complication is that auth tokens for private registries (<code>//registry.npmjs.org/:_authToken=...</code>) live in the same <code>.npmrc</code> file. None of that can leave the host. Any approach has to filter the file body before it transits anywhere.</p>
<h2 id="firstapproachmonitoringnpmrcwiththecustomlogsfilestreamintegration">First approach: Monitoring .npmrc with the Custom Logs Filestream integration</h2>
<p>The <a href="https://www.elastic.co/docs/reference/integrations/filestream">Custom Logs Filestream integration</a> in <a href="https://www.elastic.co/docs/reference/fleet">Fleet</a> wraps Filebeat's modern <code>filestream</code> input and lets you specify path globs and an ingest pipeline. Note: this is a distinct Fleet integration from the older "Custom Logs" integration, which wraps the legacy <code>log</code> input. The two have different configuration surfaces and file-identity models, so if you're migrating between them the reference linked above is worth reading before either direction.</p>
<p>We configured the integration across all users on macOS, plus the system-global npm paths:</p>
<pre><code>/Users/*/.npmrc
/opt/homebrew/etc/npmrc
/usr/local/etc/npmrc
/etc/npmrc
</code></pre>
<p>The ingest pipeline did four things:</p>
<ol>
<li><a href="https://www.elastic.co/docs/reference/enrich-processor/grok-processor">Grok</a> the <code>message</code> field to pull out the <code>min-release-age=value</code> pair.  </li>
<li>Drop any line whose key wasn't <code>min-release-age</code>, so registry auth tokens never made it into Elasticsearch.  </li>
<li>Convert the value to a <code>long</code>.  </li>
<li>Remove the raw <code>message</code> and <code>event.original</code> fields before indexing.</li>
</ol>
<p>It worked, with three caveats that changed our minds about whether it was the right integration for this job.</p>
<h3 id="threefilestreambehavioursthatdontfitconfigfilemonitoring">Three filestream behaviours that don't fit config file monitoring</h3>
<h4 id="nativefileidentitybeatsfingerprintforsmallfiles">Native file identity beats fingerprint for small files</h4>
<p>Filebeat's default file identity strategy is "fingerprint", which hashes the first N bytes of a file to give it a stable ID across rotations. The fingerprint length defaults to 1024 bytes, and a file shorter than that minimum is silently held back from ingestion until it grows. A 22-byte <code>.npmrc</code> will never grow, so the data never moves. This is in line with the description of the Filestream integration and also apparent in the agent logs:</p>
<pre><code>"ingestion from some files will be delayed, files need to be at least
 1024 in size for ingestion to start"
</code></pre>
<p>The fix is to switch the integration's advanced options to use native file identity, which uses inode plus device and has no size floor. In our deployment on Elastic Agent 9.x, disabling fingerprint alone did not fall back to native; we had to flip both toggles explicitly.</p>
<h4 id="clean_inactiveandignore_olderarecoupled">clean_inactive and ignore_older are coupled</h4>
<p>We wanted periodic re-emission of each config file so a stale event could be distinguished from a current one. filestream's <code>clean_inactive</code> is the right knob for that, but it requires <code>ignore_older</code> to also be set. Together they cleanly trim files that haven't moved recently, which is useful for rotating log files and not useful for static config files that may sit unchanged for months. We disabled <code>clean_inactive</code>.</p>
<h4 id="filestreamemitsonappendnotonremoval">Filestream emits on append, not on removal</h4>
<p>This third learning is structural rather than a config gotcha. filestream is designed for log tailing: when bytes arrive at the end of a file, an event is emitted. When content disappears (because npm rewrites <code>.npmrc</code> without the cooldown line, or because the file is deleted), filestream sees a modification, but every resulting line is dropped by our allowlist filter, and nothing reaches Elasticsearch. The last "set" event for that host stays in the index indefinitely.</p>
<p>For application logs that's correct behavior, since log lines aren't usually retracted. For config file state monitoring, removal is the signal the adoption campaign needs most, and it never arrives. That's why we switched to CEL.</p>
<h3 id="whynvminflatesnpmversioncountsforcooldownadoption">Why nvm inflates npm version counts for cooldown adoption</h3>
<p>Machines routinely report more than one npm version, because every Node.js version installed through nvm (Node Version Manager) ships its own npm binary. The distortion stood out as soon as we checked test pulls from our <a href="https://www.elastic.co/docs/reference/integrations/osquery_manager">osquery manager integration</a> package-version inventory: one machine reported 34 distinct npm versions.</p>
<p>The config side of this is forgiving. <code>npm config set min-release-age 7</code> writes to the user-level <code>~/.npmrc</code>, and every npm installation on the machine reads that file, including the per-Node copies nvm installs. Enforcement is the strict side: versions older than 11.10 ignore the key. A machine carrying npm 11.12 in one nvm environment and npm 10.9 in another is only protected when the newer npm is the one running <code>npm install</code>. A count of machines with any capable npm overstates who is protected, which is one more reason to track the config file directly rather than infer adoption from version counts.</p>
<h2 id="secondapproachmonitoringnpmrcwithcelsnapshotsemantics">Second approach: Monitoring .npmrc with CEL snapshot semantics</h2>
<p>The <a href="https://www.elastic.co/docs/reference/beats/filebeat/filebeat-input-cel">CEL input</a> in Elastic Agent was designed for HTTP API polling, but it also exposes a <code>file()</code> function and a <code>dir()</code> function for filesystem access. That enables a different model from filestream. filestream tails lines as they arrive; CEL takes a snapshot of the whole file each interval. For state monitoring, the snapshot is what we want: the unit of interest is the current contents of <code>.npmrc</code> as a whole.</p>
<p>Here is the complete script we run, trimmed to npm (our production version watches the other package managers' config files with the same pattern, and carries an extra guard for non-UTF-8 file content):</p>
<pre><code>(
  (
    try(dir("/Users")).as(entries, type(entries) != type("") ?
      entries.filter(u,
        u.is_dir
        &amp;&amp; !string(u.name).startsWith(".")
        &amp;&amp; string(u.name) != "Shared"
        &amp;&amp; string(u.name) != "Guest"
      ).map(u, "/Users/" + string(u.name) + "/.npmrc")
    : [])
  ) + (
    try(dir("/home")).as(entries, type(entries) != type("") ?
      entries.filter(u,
        u.is_dir
        &amp;&amp; !string(u.name).startsWith(".")
      ).map(u, "/home/" + string(u.name) + "/.npmrc")
    : [])
  ) + (has(state.files) ? state.files : [])
).map(f,
  try(file(f)).as(content,
    type(content) == type("") ?
      {"file": f, "exists": false}
    :
      {"file": f, "body": string(content),
       "hash": content.sha256().hex(), "exists": true}
  )
).as(file_data, {
  "events": file_data.filter(fd, fd.exists).map(fd, {
    "message": fd.body,
    "file": {"path": fd.file, "hash": {"sha256": fd.hash}},
  }),
  "cursor": {"hashes": file_data.filter(fd, fd.exists)
    .map(fd, {"file": fd.file, "hash": fd.hash})},
  "url": state.url,
  "files": has(state.files) ? state.files : [],
})
</code></pre>
<p>The first block enumerates user home directories at runtime, with filtering that path globs can't express, and appends the machine-global paths from <code>state.files</code>. <code>try(dir(...))</code> and <code>try(file(...))</code> return an error string when the path doesn't exist, so each block checks the type of the result and collapses a missing directory or file to an empty result. A macOS host has no <code>/home</code>, a Linux host has no <code>/Users</code>, and the same script runs on both.</p>
<p>The initial state supplies the global paths and the poll interval is 6 hours:</p>
<pre><code>files:
  - /opt/homebrew/etc/npmrc
  - /usr/local/etc/npmrc
  - /etc/npmrc
  - /usr/lib/node_modules/npm/.npmrc
</code></pre>
<p><strong>An aside on Fleet:</strong> the integration to add in the UI is called <a href="https://www.elastic.co/docs/reference/integrations/cel">"Custom API using Common Expression Language"</a>, not "CEL". It exposes the full CEL runtime including file system access. Set <code>resource.url: file:///dev/null</code> to satisfy the required URL field without making an HTTP request, and put the initial state in the "Custom request cursor" YAML at the bottom of the form.</p>
<h3 id="howthecelintegrationevolvedemitonchangetombstonesandheartbeat">How the CEL integration evolved: emit-on-change, tombstones and heartbeat</h3>
<p>The CEL snapshot integration shown above is the third version. The path there is the useful part.</p>
<p><strong>Version one emitted only on change.</strong> Each heartbeat hashed every file and emitted an event only when the hash differs from the cursor, the state the input persists between runs. Efficient, and it produced the removal signal cleanly: npm rewrites <code>.npmrc</code> in place when a setting changes, the hash moves, the new snapshot carries no <code>min-release-age</code> line, and the ingest pipeline marks the event <code>cooldown.absent = true</code>.</p>
<p><strong>Version two added a deletion tombstone.</strong> Hash comparison can't see a file that stopped existing (<code>npm config delete min-release-age</code> removes <code>.npmrc</code> entirely when it's the only setting), because version one filtered out non-existent files before emitting. A one-block extension emitted a stub event, once, for any file that was in the cursor but no longer on disk.</p>
<p><strong>Version three replaced both with a heartbeat, because the dashboard demanded it.</strong> The adoption dashboard counts hosts whose cooldown state falls inside the selected time window. Under emit-on-change, a host that sets a cooldown emits exactly one event and then goes silent. Once that single event ages past the dashboard's window (say, <code>now-7d</code>), the host disappears from the "adopted" count even though the cooldown is still in place. We watched adoption climb during rollout and then start to erode a few days later, purely as an artifact of one-shot events aging out of the window. The telemetry was correct; the time-windowed view of it was not.</p>
<p>So the final integration re-emits each existing file's current state on every 6-hour heartbeat. Every host re-reports its cooldown posture four times a day, the windowed dashboard reflects live state, and a host that stops appearing is genuinely offline rather than merely quiet. Removal of the cooldown line still surfaces as <code>cooldown.absent = true</code> in the next snapshot, and a deleted file drops out of subsequent snapshots, so the tombstone becomes unnecessary.</p>
<p>The cost is volume: a snapshot per file per host per interval instead of one event per change. At a 6-hour cadence that is a few thousand small documents a day across the fleet, which is immaterial. At a 60-second interval it stops being immaterial: always-emitting at 60s ships the same unchanged state every minute, which we measured at roughly a thousand-fold more documents for zero added signal. If you adopt the heartbeat, set the interval in hours.</p>
<p>The detection latency for a removal is the heartbeat interval, 6 hours in our case. For an adoption campaign measured over days and weeks against a 7-day cooldown, that is not a meaningful constraint. If you need lower-latency, discrete removal events for alerting, run the emit-on-change variant instead; the snapshot model supports both.</p>
<h3 id="howtofilternpmrcauthtokensbeforetheyleavethehost">How to filter .npmrc auth tokens before they leave the host</h3>
<p>A snapshot-based integration sends the whole file body across the wire to the ingest pipeline. That's fine for config keys; it is not fine for <code>.npmrc</code> registry auth tokens. Filtering at the ingest pipeline is too late, because by then the tokens have already transited the network.</p>
<p>The fix is an agent-side <a href="https://www.elastic.co/docs/reference/beats/filebeat/processor-script">script processor</a> in the integration's Advanced options → Processors field. It keeps only <code>min-release-age</code> lines and drops everything else before the event leaves the workstation:</p>
<pre><code>- script:
    lang: javascript
    source: &gt;
      function process(event) {
        var msg = event.Get("message");
        if (msg == null) return;
        var filtered = [];
        var lines = msg.split('\n');
        for (var i = 0; i &lt; lines.length; i++) {
          var line = lines[i].trim();
          if (line.indexOf('min-release-age') === 0) {
            filtered.push(line);
          }
        }
        event.Put("message", filtered.join('\n'));
      }
</code></pre>
<p>Tokens never reach the wire. We verified by adding a test token to a <code>.npmrc</code> on a managed host, running a cycle, and confirming zero search hits for the token string in Elasticsearch.</p>
<h3 id="thenpmcooldowningestpipeline">The npm cooldown ingest pipeline</h3>
<p>The pipeline below is the npm-only variant of the one we run in production, and it is short enough to show whole. Grok extracts the key and value, a Set processor turns "no key found" into the explicit removal signal, and the raw message is removed before indexing:</p>
<pre><code>PUT _ingest/pipeline/npm-cooldown-workstation
{
  "description": "Parses min-release-age from .npmrc snapshots",
  "processors": [
    {
      "grok": {
        "field": "message",
        "patterns": [
          "(?&lt;cooldown.key&gt;min-release-age)\\s*=\\s*(?&lt;cooldown.value&gt;[^\\n\\r]*)"
        ],
        "ignore_missing": true,
        "ignore_failure": true
      }
    },
    {
      "set": {
        "if": "ctx.cooldown?.key == null",
        "field": "cooldown.absent",
        "value": true
      }
    },
    {
      "set": {
        "if": "ctx.cooldown?.key != null",
        "field": "cooldown.unit",
        "value": "days"
      }
    },
    {
      "convert": {
        "field": "cooldown.value",
        "type": "long",
        "ignore_missing": true,
        "ignore_failure": true
      }
    },
    {
      "remove": {
        "field": ["message", "event.original"],
        "ignore_missing": true
      }
    }
  ]
}
</code></pre>
<p>Two details earned their place the hard way. The value capture is <code>[^\n\r]*</code>, which stops at both LF and CRLF line endings; an earlier revision used <code>[^ \r]*</code>, which truncated any value containing a space and could run across newlines. And <code>cooldown.unit</code> is set explicitly even though npm's unit is always days, because dashboard rows that read <code>7 days</code> stay unambiguous when the telemetry later grows beyond npm.</p>
<p>The <code>cooldown.absent = true</code> event is the payoff. The host was in the index six hours ago with <code>cooldown.key = min-release-age</code>; the next snapshot has no key; the dashboard shows the transition. That is the removal signal filestream could not produce.</p>
<h2 id="celvsfilestreamforconfigfilemonitoring">CEL vs. filestream for config file monitoring</h2>
<p>We ran both in parallel for a few weeks and then made a call.</p>
<p>The core reason is the shape of the data. filestream is built for append-only logs, where new lines arrive at the end of a file and the job is to harvest them. <code>.npmrc</code> is a state file: npm rewrites it in place when a setting changes and deletes it when the last setting is removed. The question the telemetry needs to answer, "what is the cooldown state of this host right now, and when did it change," is a question about the current contents of a file. Tail offsets can't answer it.</p>
<p>|  | CEL | Filestream |
| :---- | :---- | :---- |
| Designed for | State files (snapshot + hash) | Append-only logs (tail) |
| Cooldown-line removal detection | Yes (next snapshot marks <code>cooldown.absent</code>) | No |
| File deletion detection | Yes (host drops out of snapshots) | No |
| Snapshot semantics | Whole file | Per line |
| Auth-token posture | Filtered agent-side before transit | Same (agent-side allowlist) |
| Fleet integration | Custom API using Common Expression Language | Custom Logs Filestream |
| Config complexity | ~40-line CEL integration | Declarative path globs |
| State refresh / latency | 6h heartbeat | ~10s (tail) |</p>
<p>Running both meant every config change produced two events into two separate datasets, and we maintained two ingest pipelines with dashboards split across indices. The complexity cost outweighs any redundancy benefit, and the only thing filestream gives that CEL doesn't is faster detection latency, which is not a meaningful constraint for adoption tracking on a 7-day cooldown.</p>
<p>We also considered osquery and auditd before settling on CEL. osquery can read config file contents on a schedule through its <code>file</code> and <code>file_lines</code> tables, and auditd can fire on file writes, but neither produces a clean current-state signal across ingestion, and both would mean standing up a second collection path next to the agent already deployed for endpoint telemetry.</p>
<p>CEL won. End-to-end validation on a single macOS workstation:</p>
<p>| Scenario | Expected | Result |
| :---- | :---- | :---- |
| <code>npm config set min-release-age 5</code> | <code>cooldown.key = min-release-age</code>, <code>cooldown.value = 5</code> | Pass |
| <code>.npmrc</code> exists with no cooldown key | <code>cooldown.absent = true</code> | Pass |
| <code>npm config delete min-release-age</code> (deletes the file) | path absent from subsequent snapshots | Pass |
| Auth token line in <code>.npmrc</code> | No token in Elasticsearch | Pass |</p>
<h2 id="npmcooldownmonitoringonlinuxandwindows">npm cooldown monitoring on Linux and Windows</h2>
<p>macOS was the primary platform; the design generalizes cleanly.</p>
<p>Linux needed no integration change at all. The <code>/home</code> block in the integration above already covers it: on each platform, the directory that doesn't exist fails the type check and collapses to an empty list, and the other side fills in. The Linux variant runs the same integration through the same ingest pipeline into the same dataset, with <code>/etc/npmrc</code> and <code>/usr/lib/node_modules/npm/.npmrc</code> in the global watch list.</p>
<p>Windows is a separate integration on the same dataset, and it's the one variant we designed but haven't deployed: our own fleet has too few Windows machines to justify the rollout yet. We're documenting it anyway because Windows is the most common OS in corporate fleets, so for many readers this variant is the one that matters most. The integration enumerates <code>C:/Users/*</code> (forward slashes work on Windows under Go/CEL), excludes <code>Public</code>, <code>Default</code>, and <code>Default User</code>, and watches each user's <code>.npmrc</code> plus the system globals under <code>C:/ProgramData</code>. The Grok pattern's value capture, <code>[^\n\r]*</code>, stops at both LF and CRLF line endings, so it works unchanged on Windows without capturing the trailing <code>\r</code> in the value. The agent-side script processor is also unchanged.</p>
<h2 id="rolloutstatusandextendingbeyondnpm">Rollout status and extending beyond npm</h2>
<p>The pipeline is now reporting from several hundred macOS, each re-emitting its current npm cooldown state on the 6-hour heartbeat. We're sharing the technique at this stage so other security engineering teams can build on it; once the rollout reaches the full fleet we'll follow up with an org-wide adoption-rate post. The Windows variant stays on the shelf until our Windows population justifies the rollout; if your fleet is Windows-heavy, the design in this post is ready to adapt.</p>
<p>In production the same integration already watches the config files for pip, uv, pnpm, yarn Berry, and bun. That expansion deserves its own post, because cooldown values are not unit-comparable across package managers, and one manager's documentation and implementation disagree about the unit by three orders of magnitude. If you're extending this design beyond npm, check <a href="https://cooldowns.dev">cooldowns.dev</a> for each flag's unit before you compare values across tools.</p>
<h2 id="whatwelearnedaboutnpmcooldownmonitoringwithelasticagent">What we learned about npm cooldown monitoring with Elastic Agent</h2>
<ul>
<li>npm cooldown adoption telemetry is state monitoring. The signal that matters most is a host removing <code>min-release-age</code>, and append-driven log tailing cannot see that happen.  </li>
<li>A snapshot-based CEL integration produces the removal signal directly: re-emit each <code>.npmrc</code>'s current state on a 6-hour heartbeat, and a removal shows up as the key's absence in the next snapshot.  </li>
<li>A time-windowed adoption dashboard needs the heartbeat, not one-shot change events, which age out of the window and make adoption look like it's eroding when it isn't.  </li>
<li>Small config files (under 1024 bytes) need native file identity, not fingerprint, regardless of which integration you use.  </li>
<li>An agent-side script processor keeps <code>.npmrc</code> registry auth tokens on the workstation: tokens never reach the wire.  </li>
<li>nvm inflates capability counts: each Node.js version ships its own npm, all of them read the shared <code>~/.npmrc</code>, but only npm 11.10+ enforces the key. Track the config file, count enforcement-capable versions separately.</li>
</ul>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/npm-cooldown-removal-detection-elastic-agent</link>
    <guid isPermaLink="false">npm-cooldown-removal-detection-elastic-agent</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Wieger van der Meulen]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbd7f6e8ecebcfe29/6a7d8394a529e16d0759c958/cover.png" length="0" type="image/png"/>
    <pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elastic goes all-in on Hacker Summer Camp at Black Hat and DEF CON in Las Vegas]]></title>
    <description><![CDATA[Attack Discovery turns raw alerts into validated threats and Elastic Defend closes vulnerable driver gaps as fast as they're disclosed. Watch it all run against real attacks at the booth.]]></description>
    <content:encoded><![CDATA[<p>At Elastic, we know that the best way to build security tools is to bring them to the community, have security pros use them, and let them tell us what features and functionality matter and why. This year, we’re excited to do this at Black Hat and DEFCON, the weeklong security marathon affectionately known as Hacker Summer Camp. Our smartest technical experts and practitioners will be at Black Hat showing off our latest innovations, sponsoring and hosting events to help security experts and leaders connect, and at DEFCON’s Blue Team Village with our new Capture the Flag challenge to help defenders sharpen their investigation skills.</p>
<p>This community-powered innovation is evident in everything we do. Elastic is a security tool built by security users, for security users. We’ve sat in the seat. We’ve worked the queue at 2 a.m. We’ve chased an alert that turned out to be nothing and missed the one that turned out to be everything. What we're continually building and improving is the security operations center (SOC) we wished we'd had back then. This means agents that carry the machine-speed work, leave critical judgment to analysts, and a platform that connects the two.</p>
<p>With Elastic, machine speed and human judgment work together in a single loop. Stopping more at the endpoint reduces the number of alerts. Those that remain surface the real threats, and you can validate them before they reach a queue. The work underneath is increasingly automated. Each piece makes the next one lighter, and none of it asks you to hand judgment over to a black box.</p>
<h2 id="alertzerofromalertqueuetovalidatedthreats">Alert Zero: From alert queue to validated threats</h2>
<p>Every SOC is chasing a queue worked down to what actually matters, the SOC's version of “inbox zero.” When we built our suite of tools, our goal was <a href="https://www.elastic.co/security-labs/agentic-soc-alert-triage-alertzero">Alert Zero</a>, a state that always felt out of reach. It’s a goal that teams move toward, with agents and analysts working together. It doesn’t mean zero alerts or replacing the analysts.</p>
<h3 id="howattackdiscoveryinvestigatesalertslikeananalyst">How Attack Discovery investigates alerts like an analyst</h3>
<p>Attack Discovery has always pulled related alerts together into a single view of an attack. Now it goes further, working through them the way a human analyst would:</p>
<ol>
<li>Threat-hunts raw events beyond the initial alerts.  </li>
<li>Checks entity risk for the users and hosts involved.  </li>
<li>Corroborates findings across other data sources.  </li>
<li>Classifies the event as a validated attack.</li>
</ol>
<p>Your team gets a short list of validated attacks to work, instead of a wall of raw alerts to triage.</p>
<h3 id="closingdetectiongapswithautodraftedrules">Closing detection gaps with auto-drafted rules</h3>
<p>When Attack Discovery finds something that your rules missed, it drafts a detection rule to close the gap and hands it to an analyst to approve, helping to make the entire workflow more efficient and to reduce the source of false positives. </p>
<p>Security teams need the <em>how,</em> not just the <em>what</em>, and Attack Discovery shows its work, so you can see how it got to each answer and recommended action. Every step of the reasoning is visible, so an analyst knows why an alert became an attack. You can run it however fits your team, whether you kick it off yourself or set a recurring cadence. You can even trigger it from  Elastic Workflows. A separate alert analysis workflow addresses the volume from the other side, differentiating between likely false and true positives, so analysts lose fewer hours to low-fidelity alerts, and leaving Attack Discovery a cleaner set to investigate.</p>
<h2 id="elasticdefendendpointprotectionvulnerabledrivercoverageandwindowsonarm">Elastic Defend endpoint protection: vulnerable driver coverage and Windows on ARM</h2>
<p>Fewer alerts reach the queue when more threats are stopped on the device, so prevention starts at the endpoint.</p>
<h3 id="vulnerabledrivercoveragethatkeepspacewithdisclosure">Vulnerable driver coverage that keeps pace with disclosure</h3>
<p>Elastic Defend <a href="https://www.elastic.co/security-labs/vulnerable-driver-detection-elastic-defend-byovd">now gets ahead of vulnerable drivers</a>. Attackers exploit these by bringing a signed, trusted driver with a known flaw and using it to reach the kernel, and coverage for a new one has traditionally arrived on a release cycle. Our <a href="https://www.elastic.co/security-labs">threat research team</a> monitors public disclosure sources, like VirusTotal, loldrivers.io, and Microsoft's blocklist. Through an always-on process, Elastic automatically generates and instantly deploys YARA rules as new drivers are disclosed, so protection keeps pace instead of waiting on a release. That speed matters when AI-driven attacks can move from one machine to the next in under a minute, faster than any response workflow can react.</p>
<h3 id="fullendpointprotectionforwindowsonarm">Full endpoint protection for Windows on ARM</h3>
<p>Windows on ARM is now fully covered in Defend, which brings Surface and other ARM-based laptops into the same protection as the rest of your fleet. Teams can roll this feature out across every endpoint without paying per device to do so. A new endpoint troubleshooting skill also rounds out this capability, automatically flagging policy and performance issues so your team spends less time chasing them.</p>
<h2 id="elasticworkflowssocautomationyoucandescribeinplainlanguage">Elastic Workflows: SOC automation you can describe in plain language</h2>
<h3 id="buildautomationsinplainlanguagewithversioncontrol">Build automations in plain language with version control</h3>
<p>Elastic Workflows makes automations faster to build and shows you exactly what a workflow will do before it runs. The latest updates start with <a href="https://www.elastic.co/search-labs/blog/ai-workflow-automation-natural-language">plain-language authoring</a>, so you can describe the automation you want and have it generated for you. Versioning then tracks every change, so you can compare any two versions and roll back to a working one in a click. You always know who changed what and when. Visual Mode shows a workflow as a graph, with its triggers, steps, branches, and logic visible at a glance next to the YAML. Drag-and-drop editing is coming next.</p>
<h3 id="humanintheloopapprovalsroutedtoslack">Human-in-the-loop approvals routed to Slack</h3>
<p>Automation you can trust starts with understanding what Workflows will do and when it will ask for help. When a workflow reaches a decision that needs a person or an approval or other input, it pauses and routes the request to a tool your team already uses, like Slack. Automation handles the routine, while your team stays on top of the decisions that need judgment. Workflows runs natively within the Elasticsearch platform extending across search, observability, and security, so it runs where your security data already lives, rather than stitched as a layer on top.</p>
<p>Together, our drive to Alert Zero, enhanced endpoint protection, and automation where your data lives reinforce each other. Stronger prevention keeps alerts from being raised in the first place, and the ones that remain arrive validated instead of raw. Automation underneath keeps prevention and investigation moving at machine speed, while your analysts stay on the decisions that need a human. That’s what the agentic SOC looks like when it’s built to help the people in it rather than replace them.</p>
<p><em>"Security teams are not losing because they lack tools; they're losing because the tools generate more work than the team can absorb," said <strong>Mike Nichols, general manager Security, Elastic.</strong> "Elastic Security is built by people who've sat in the SOC and worked the queue. These updates go after one of the biggest sources of analyst burnout, which are alerts that shouldn't be alerts in the first place. We know every barrier is a liability and we’re building an open, transparent platform that breaks down those barriers and makes it easier for teams to customize and manage the security stack they need to protect their organizations."</em></p>
<h2 id="findelasticsecurityatblackhatanddefcon2026">Find Elastic Security at Black Hat and DEF CON 2026</h2>
<p>At the booth, you can see all of this working on real attacks:</p>
<ul>
<li>Agents helping cut false positives on the way to Alert Zero, showing their reasoning at every step.  </li>
<li>Security information and event management (SIEM) built for the agentic SOC.  </li>
<li>Extended detection and response (XDR) across cloud, Kubernetes, and endpoint.  </li>
<li>Native automation with Elastic Workflows.  </li>
<li>Endpoint defense that holds across Linux, macOS, and Windows.  </li>
<li>Threat research from Elastic Security Labs that ships in the platform.</li>
</ul>
<p>Bring your hardest questions.</p>
<p>Elastic is showing up across the week, at Black Hat and around the community.</p>
<ul>
<li><strong>Sober Speakeasy</strong> with Sober in Cyber, an alcohol-free networking evening at the Mob Museum's Underground Speakeasy. Tuesday, August 4, 2026, 7:00–9:30 p.m. PT. All infosec professionals are welcome, sober or sober-curious.  </li>
<li><strong>Blitz &amp; Defend</strong>, an executive evening with Elastic and AWS at Allegiant Stadium, a behind-the-scenes tour, and reception in the Raiders locker room. Wednesday, August 5, 2026, 6:30-8:30 p.m. PT.  </li>
<li><strong>GDIT Sip &amp; Cipher</strong>, a cybersecurity networking reception at Toca Madera, with Elastic among the sponsors. Wednesday, August 5, 2026, 6:00–9:00 p.m. PT.  </li>
<li><strong>InnovatHERs: Women Shaping Tomorrow</strong>, a breakfast with Women in CyberSecurity (WiCyS) to celebrate and connect women in the field. Thursday, August 6, 2026, 7:00–9:00 a.m. PT.  </li>
<li><strong>Find us at DEF CON, too:</strong> Security people go to Black Hat because they have to. They go to DEF CON because they want to. That's why we're proud to be a Blue Tier sponsor of Blue Team Village at DEF CON this year, the highest-traffic village and home to the SOC, Digital Forensics and Incident Response (DFIR), and incident response communities. It's where defenders come to sharpen their craft, run capture-the-flag, and swap real-world detection and response tactics. For us, showing up there is about being where our people already are, because Elastic Security is built by the same community that fills the room. Come find us.</li>
</ul>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/elastic-security-black-hat-defcon-2026</link>
    <guid isPermaLink="false">elastic-security-black-hat-defcon-2026</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Jackie McGuire]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt281ebbe438de2876/6a7d7f7fbd21984bcd75527f/cover.png" length="0" type="image/png"/>
    <pubDate>Fri, 31 Jul 2026 21:59:59 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elastic on Defence Cyber Marvel 2026: A Technical overview from the Exercise Floor]]></title>
    <description><![CDATA[An overview of the Elastic Security and AI infrastructure deployed to support the UK Ministry of Defence's flagship cyber exercise, Defence Cyber Marvel 2026.]]></description>
    <content:encoded><![CDATA[<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt50d0090f9fe0c606/6a9fe114de23957070e85c31/DCM2026_-_Logos_(1).png" alt="enter image description here" /></p>
<p>Where to begin. For the fourth consecutive year, Elastic has had the privilege of serving as a trusted industry partner on Exercise Defence Cyber Marvel - the UK Ministry of Defence's flagship cyber exercise series. DCM26 was, without question, the most ambitious iteration yet, and we're chuffed to bits to finally be able to talk about what we built, how we built it, and what we learnt along the way.</p>
<h2 id="whatisdefencecybermarvel">What is Defence Cyber Marvel?</h2>
<p>For those unfamiliar, Defence Cyber Marvel (DCM) is the largest UK military cyber exercise series that focuses on defending traditional IT networks, corporate environments, and complex industrial control systems in realistic, high-pressure scenarios. It showcases responsible cyber power whilst enhancing readiness, interoperability, and resilience across Defence and allied nations. Now in its fifth year, DCM has evolved from an Army Cyber Association initiative into a tri-service operation led by Cyber and Specialist Operations Command (CSOC).</p>
<p>The <a href="https://www.gov.uk/government/news/uk-to-lead-multinational-cyber-defence-exercise-from-singapore">UK Government published an official press release for DCM26</a>, which provides an excellent overview of the exercise's strategic importance. As the British High Commissioner to Singapore noted, the exercise demonstrates the deep cooperation between the UK and trusted partners, a reminder of the strength of shared strategic partnerships in an increasingly complex security landscape.</p>
<p>At its core, DCM is a force-on-force cyber exercise: defending Blue Teams protect their assigned networks and infrastructure from attacking Red Teams, using a range of techniques. Activities span changing default passwords and hardening firewalls through to deploying enterprise-grade, AI-powered cyber defence with <a href="https://www.elastic.co/security">Elastic Security</a>. The activities of each team are monitored by the White Team to establish a score factoring in system availability, attack detection, incident reporting, and system restoration.. It stretches the most experienced teams whilst also facilitating a unique training mechanism for junior teams on their first exposure to a cyber range, and that dual purpose is what makes DCM such a valuable exercise.</p>
<h2 id="thescaleofdcm26">The scale of DCM26</h2>
<p>DCM26 brought together over 2,500 personnel from 29 participating countries and 70 organisations, coordinated from a central Exercise Control (EXCON) based out of Singapore, with EXCON hosting over 600 participants. The exercise ran across a hybrid compute environment spanning the CR14 cyber range and AWS, hosting over 5,000 virtual systems.</p>
<p>The exercise itself ran for five days of execution (9–13 February 2026), preceded by optional instructor-led pre-training and connectivity checks. The scenario, built on the Defence Academy Training Environment (DATE) Indo-Pacific Operating Environment, placed teams as Cyber Protection Teams defending deployed military systems during an escalating regional crisis.Blue Teams were geographically dispersed,some in their home locations across the UK and internationally, others deployed overseas, all connecting into the range via VPN. </p>
<p>Participants included representatives from UK Defence, cross-government departments such as the National Crime Agency, the Department for Work and Pensions, the Cabinet Office, and the Department for Business and Trade, alongside international partners forming up to 40 teams. Following the success of last year's exercise in the Republic of Korea, Singapore served as the exercise hub for the first time, reflecting the UK's commitment to deepening cooperation with Indo-Pacific partners on shared security challenges.</p>
<p>In short, it's a serious exercise. High-pressure, force-on-force, with real consequences for scoring and real learning outcomes for every participant.</p>
<h2 id="thedeploymentsourelasticinfrastructure">The deployments: Our Elastic infrastructure</h2>
<p>This year's infrastructure represented a significant architectural evolution from previous iterations. Rather than deploying individual Elastic Cloud clusters per team, we moved to a single, space-based multi-tenanted Elastic Cloud deployment for the Blue Teams. We also provided deployments for functions outside of  the Blue Teams. Let me break down each deployment and why it exists.</p>
<h3 id="blueteamsmultitenantedelasticsecurity">Blue Teams: Multi-tenanted Elastic Security</h3>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta3cb5f15b25a8b42/6a7d7f648fc2d012f43eb849/image4.png" alt="" /></p>
<p>The centrepiece of our contribution was a single Elastic Cloud deployment serving all 40 defending Blue Teams, separated using Kibana Spaces and datastream namespaces. Each of the 39 teams had its own isolated workspace, including dashboards, agents, and detection rules.</p>
<p>Here's what the Terraform resource looked like for creating each team's space:</p>
<pre><code># Create 40 Blue Team spaces
resource "elasticstack_kibana_space" "blue_team" {
  count = var.team_count

  space_id    = local.space_ids[count.index]
  name        = "Blue Team ${local.team_numbers[count.index]}"
  description = "Isolated space for BT-${local.team_numbers[count.index]} with space-aware Fleet visibility"

  disabled_features = []
  color             = "#0077CC"
}
</code></pre>
<p>Each team's space got a dedicated set of three  <a href="https://www.elastic.co/docs/reference/fleet/agent-policy">Fleet</a> agent policies: on day 1a Deployed network policy, day 2, a Host Nation network policy, and finally a PacketCapture policy for network traffic monitoring. The phased access control was elegant in its simplicity: setting <code>enable_hostnation_network = true</code> in our <code>terraform.tfvars</code> and running <code>terraform apply</code> expanded each team's role permissions and made their Host Nation agent policy visible in their space. The exercise went from one network to two without a single manual click in Kibana.</p>
<p>The data isolation relied on datastream namespaces. Each agent policy is written to team-specific namespaces like <code>bt_01_deployed</code> and <code>bt_01_hostnation</code>, producing data streams following the pattern:</p>
<pre><code>logs-system.auth-bt_01_hostnation
logs-system.syslog-bt_01_hostnation
metrics-system.cpu-bt_01_hostnation
logs-endpoint.events.process-bt_01_hostnation
logs-windows.forwarded-bt_01_hostnation
logs-auditd.log-bt_01_hostnation
</code></pre>
<p>Each team's Kibana security role was then scoped to only those data streams using dynamic index privilege blocks:</p>
<pre><code># Deployed data streams (always granted)
indices {
  names = [
    "logs-*-${local.deployed_namespaces[count.index]}",
    "metrics-*-${local.deployed_namespaces[count.index]}",
    ".fleet-*"
  ]
  privileges = ["read", "view_index_metadata"]
}

# HostNation data streams (conditional on enable_hostnation_network)
dynamic "indices" {
  for_each = var.enable_hostnation_network ? [1] : []
  content {
    names = [
      "logs-*-${local.hostnation_namespaces[count.index]}",
      "metrics-*-${local.hostnation_namespaces[count.index]}"
    ]
    privileges = ["read", "view_index_metadata"]
  }
}
</code></pre>
<p>Authentication was handled via Keycloak SSO, with Elasticsearch role mappings connecting Keycloak groups to Kibana roles:</p>
<pre><code>resource "elasticstack_elasticsearch_security_role_mapping" "blue_team" {
  count = var.team_count

  name    = "bt-${local.team_numbers[count.index]}-keycloak-mapping"
  enabled = true

  roles = [
    elasticstack_kibana_security_role.blue_team[count.index].name
  ]

  rules = jsonencode({
    field = {
      groups = "${local.keycloak_groups[count.index]}"
    }
  })
}
</code></pre>
<p>The default integration policies were simple by design. Each team received: System for core OS telemetry, Elastic Defend for Endpoint Detection and Response, Windows event forwarding, Auditd for Linux audit logging, and Network Packet Capture integrations. That's over 400 integration policies managed as code via the <a href="https://registry.terraform.io/providers/elastic/elasticstack/latest/docs">Elastic Stack Terraform Provider</a>.</p>
<p>A note on Elastic Defend: due to the effectiveness of Elastic's endpoint protection - which is trusted in production by the <a href="https://www.elastic.co/blog/defense-and-intelligence-community-endpoint-security">US DOD and IC, read more about that here</a> - and the fact that nobody in their right mind is burning zero-day exploits on a training exercise, we're forced tohandicap Elastic Defend by disabling Prevent mode, leaving it in Detect-only mode. Teams get alerts when something malicious happens, but with no automatic mitigation. We also completely disable Memory Threat Prevention and Detection as this discovers the majority of attacking team implants and beacons, which would rather spoil the game for the Red Teams. Toward the end of the exercise, we allowed the teams the freedom to use Elastic Defend to its full capability, but not before letting the Red Teams get a strong foothold.</p>
<p>We also pre-installed Elastic's <a href="https://www.elastic.co/docs/reference/security/prebuilt-rules">prebuilt detection rules</a> into each team space - the full set from Elastic Security Labs, continuously updated in an open repository. These rules were setup to ensure they only queried indices that the team's namespace-scoped permissions allowed, preventing any cross-team data leakage in detection rule execution.</p>
<p>Additionally, each team space had its Security Solution default index configured to scope detection rules to only that team's data streams, rather than the default broad pattern. This was handled by a Terraform <code>null_resource</code> that called the Kibana internal settings API to set <code>securitySolution:defaultIndex</code> for each space.</p>
<p>At peak, this deployment was ingesting 800,000 events per second (EPS) across all 40 teams. That's a serious amount of data, and the cluster handled it comfortably thanks to the autoscaling capabilities of Elastic Cloud. <a href="https://www.elastic.co/blog/monitoring-petabytes-of-logs-at-ebay-with-beats">That given, back in 2018 we were doing 5 million events per second with eBay.</a></p>
<p>Data lifecycle was managed by an Index Lifecycle Management (ILM) policy that rolled indices over after one day or <code>50</code> GB (whichever came first), moved them to a warm phase after two days for read-only optimisation and force-merging, and then deleted data after ten days. As a result, the storage costs were minimized while maintaining the exercise window requirements. Below is an example of how the ILM policy was implemented.</p>
<pre><code>resource "elasticstack_elasticsearch_index_lifecycle" "dcm5_10day_retention" {
  name = "dcm5-10day-retention"

  hot {
    min_age = "0ms"

    set_priority {
      priority = 100
    }

    rollover {
      max_age                = "1d"
      max_primary_shard_size = "50gb"
    }
  }

  warm {
    min_age = "2d"

    set_priority {
      priority = 50
    }

    readonly {}

    forcemerge {
      max_num_segments = 1
    }
  }

  delete {
    min_age = "${var.data_retention_days}d"

    delete {
      delete_searchable_snapshot = true
    }
  }
}
</code></pre>
<h3 id="theshardstresstestprovingmultitenancyatscale">The shard stress test: Proving multi-tenancy at scale</h3>
<p>Before committing to this architecture for a live military exercise, we needed to prove it would be able to meet our requirements and have an appropriate failover in place in the event of issues. Moving from individual deployments to a single multi-tenanted cluster introduced real risks: resource contention, ingest bottlenecks, data leakage across spaces due to misconfiguration, large TCP connection counts on the Elasticsearch nodes, and a significantly larger shard count since each team generates its own set of indices.</p>
<p>So we built a dedicated testing rig. The plan was straightforward: deploy 50 Kibana Spaces, create an agent policy in each space, launch 6,000 EC2 instances (120 per tenant, across six subnets in three availability zones), and load-test the lot. We monitored everything with AutoOps and Stack Monitoring.</p>
<p>The deployment flow worked like this: Terraform created the VPC and subnets across three availability zones, provisioned the 50 Kibana Spaces and their space-scoped Fleet policies, generated enrolment tokens, and then launched EC2 instances in batches. Each instance installed Elastic Agent on boot and enrolled against its space-specific token.</p>
<p>We hit some interesting challenges along the way. The standard Elastic Stack Terraform Provider didn't support space-aware Fleet operations at the time, so we forked it and added space ID handling to the Fleet resources - without that modification, every agent would have enrolled into the default space regardless of policy assignment. This wasn't the first time we'd had to extend the provider for an exercise; two years ago, for DCM2, we'd added the <code>elasticsearch_cluster_info</code> data source. Fortunately, the upstream provider has since added <code>support for space_ids</code> in version <code>0.12.2</code>.</p>
<p>We also ran into AWS EC2 API rate limits when trying to spin up all 6,000 instances simultaneously, so we batched deployments at 500 instances with five-minute cool-off periods between batches.</p>
<p>The results were reassuring. All 6,000 agents were typically enrolled within 20 minutes of deployment. In our tests, space isolation worked as expected with no observed data leakage between tenants. Fleet policy updates propagated to all agents within 60 seconds. Search queries scoped to individual spaces remained fast under full load. And the multi-AZ distribution proved resilient during simulated availability zone failures.</p>
<p>This testing gave us the confidence to commit to the architecture for the live exercise.</p>
<h3 id="redteamsc2implantobservability">Red Teams: C2 implant observability</h3>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltba4680e131e12972/6a7d7f663ce8e2bf1dcf2628/image3.png" alt="" /></p>
<p>A separate, dedicated Elastic deployment was stood up for the Red Teams, focused on Command and Control (C2) implant observability. This gave the attacking teams visibility into their own operations, including implant status, beacon callbacks, and operational progress, without any risk of cross-pollination with the Blue Team's data. The Red Teams used Tuoni as their C2, which is a framework developed by Clarified Security for red teaming. In DCM3, we worked with Clarified Security to ensure it properly supported the Elastic Common Schema, making future integration with Elastic much easier.</p>
<h3 id="nsocexercisenetworksecurityoperationscentre">NSOC: Exercise Network Security Operations Centre</h3>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt958bc861d15113ad/6a7d7f6add26d2eea42a722f/image6.png" alt="" /></p>
<p>The core exercise, Network Security Operations Centre (NSOC), ran on its own Elastic deployment, providing the exercise control staff with an overarching view of range health, security monitoring across the entire infrastructure, and critically, audit logging for all the AI services we deployed. Every <a href="https://www.elastic.co/docs/reference/integrations/aws_bedrock">Bedrock API invocation was logged in CloudWatch</a> and observable in this deployment, meaning the NSOC had complete visibility into what was being asked to the AI agents and by whom . More on this in the AI section below.</p>
<h2 id="infrastructureautomationterraformandcatapult">Infrastructure automation: Terraform and Catapult</h2>
<p>Everything you've seen above was managed as Infrastructure as Code. Our <code>provider.tf</code> gives a sense of the provider ecosystem we were orchestrating:</p>
<pre><code>terraform {
  required_version = "&gt;= 1.5"

  required_providers {
    elasticstack = {
      source  = "elastic/elasticstack"
      version = "~&gt; 0.13.1"
    }
    aws = {
      source  = "hashicorp/aws"
      version = "~&gt; 5.0"
    }
    vault = {
      source  = "hashicorp/vault"
      version = "~&gt; 3.20"
    }
    cloudflare = {
      source  = "cloudflare/cloudflare"
      version = "~&gt; 5.15.0"
    }
  }

  backend "s3" {
    bucket  = "elastic-terraform-state-dcm5"
    key     = "prod/terraform.tfstate"
    region  = "eu-west-2"
    encrypt = true
  }
}
</code></pre>
<p>The total resource footprint managed by Terraform was substantial: one Elastic Cloud deployment with autoscaling, 40 Kibana Spaces, 120 Fleet agent policies (three per team), 400+ integration policies, 40 Kibana security roles, 40 Keycloak role mappings, ILM policies for data retention, 41 AWS IAM users for Bedrock GenAI connectors (one per team space plus a default), 41 Kibana GenAI action connectors, AWS Bedrock guardrails, Cloudflare Zero Trust tunnels for Tines access, Tines action connectors per team space, detection service accounts stored in HashiCorp Vault, and per-space Security Solution default index configuration. All state was stored in an encrypted S3 backend.</p>
<p>For the agent and proxy deployment onto the actual range systems, we used <a href="https://github.com/ClarifiedSecurity/catapult">Catapult</a>, an excellent open-source tool built by the team at Clarified Security. Catapult wraps Ansible with a container-based execution model that's purpose-built for cyber range deployments. It handled the installation and enrolment of Elastic Agents across the range infrastructure. The configuration of proxy servers (each team had a dedicated Squid proxy for its deployed network, this was to simulate a single point of egress as it would be in the real world. Traffic was routing through endpoints like <code>http://elastic-proxy.dsoc.XX.dcm.ex:3128</code>), and the deployment of Cloudflare tunnels for Tines connectivity.</p>
<p>During provisioning, the following were written to  HashiCorp Vault by Terraform and consumed by Catapult: Credentials, enrolment tokens, API keys, proxy configurations, Tines service account credentials.. The Vault paths followed a consistent structure like <code>dcm/gt/elastic/prod/enrollment_tokens/BT-XX-Deployed</code> and <code>dcm/gt/elastic/tines-sa/tines-sa-btXX</code>, making it straightforward for the Catapult playbooks to pull the right credentials for each team.</p>
<h2 id="trainingsettingteamsupforsuccess">Training: setting teams up for success</h2>
<p>Deploying the platform is one thing; ensuring people can actually use it is another. We provided on-range, instructor-led training to the Blue Teams during the pre-exercise phase. This covered <a href="https://github.com/ClarifiedSecurity/catapult">Elastic Security</a> fundamentals, navigating their team space in Kibana, working with the prebuilt detection rules, using Discover for log analysis and threat hunting, building custom dashboards, understanding Elastic Defend alerts, and getting familiar with the Timeline investigation tool.</p>
<p>The exercise instruction itself noted this training was optional but "highly recommended," and from what we saw, the teams who attended absolutely hit the ground running on Day one of execution. Training and enablement are just as important as the technology deployment itself. Handing a team enterprise-grade security tooling which they don't know how to use would'nt have been helpful for anyone.</p>
<h2 id="theonrangeaiservicecompliantauditedguardrailed">The On-Range AI service: Compliant, audited, Guardrailed</h2>
<p>This year marked our debut in providing AI access to the DCM range. We provided a compliant AI service directly on the range, backed by UK-tenanted AWS Bedrock models - specifically Claude 3.7 Sonnet running in the eu-west-2 (London) region. This wasn't AI for the sake of AI; it was a carefully architected service with guardrails, complete audit logging, and RBAC-aware access controls. We were trusted with running this service due to Elastic's experience in the AI space. </p>
<p>The AI service had multiple consumers on the range, and this is an important distinction. The compliant Bedrock connector we provisioned into each team's space wasn't just powering our custom agents - it also powered Elastic's native AI features, specifically:</p>
<h3 id="elasticaiassistantforsecurity">Elastic AI Assistant for Security</h3>
<p>The <a href="https://www.elastic.co/docs/solutions/security/ai/ai-assistant">Elastic AI Assistant</a> was available in every Blue Team space, connected to our on-range Bedrock connector. This gave teams a context-aware chat interface directly within Elastic Security where they could ask questions about their alerts, get help writing ES|QL queries, investigate suspicious processes, and get guided remediation steps. The AI Assistant uses Retrieval-Augmented Generation (RAG) with Elastic's Knowledge Base feature, which is pre-populated with articles from <a href="https://www.elastic.co/security-labs">Elastic Security Labs</a>. Teams could also add their own documents, such as range-specific SOPs, threat intel, or team notes, to the Knowledge Base to further ground the assistant's responses in their operational context.</p>
<p>What made this particularly valuable in the exercise context was the AI Assistant's ability to help less experienced analysts understand what they were looking at. A junior analyst facing their first live implant beacon could ask the assistant to explain the alert, suggest investigation steps, and even help draft the incident report. The data anonymisation settings ensured that sensitive field values could be obfuscated before being sent to the LLM provider.</p>
<h3 id="elasticattackdiscovery">Elastic Attack Discovery</h3>
<p><a href="https://www.elastic.co/docs/solutions/security/ai/attack-discovery">Attack Discovery</a> was another significant consumer of our on-range AI service. Attack Discovery uses LLMs to analyse alerts in a team's environment and identify threats by correlating alerts, behaviours, and attack paths. Each "discovery" represents a potential attack and describes relationships among multiple alerts - telling teams which users and hosts are involved, how alerts map to the <a href="https://www.elastic.co/docs/solutions/security/detect-and-alert/mitre-attack-coverage">MITRE ATT&amp;CK matrix</a>, and which threat actor might be responsible.</p>
<p>For a cyber exercise in which Red Teams actively launched coordinated attacks, Attack Discovery was transformative. Instead of manually triaging hundreds of individual alerts, Blue Teams could run Attack Discovery to surface the high-level attack narratives, for example, "these 15 alerts are all part of a lateral movement chain from host X to host Y, likely by threat actor Z", and focus their investigation time where it mattered most. It's the kind of capability that directly reduces mean time to respond, and fights alert fatigue, which is precisely what you need when you're under sustained attack for five days straight.</p>
<h2 id="thecustomaiagentselasticagentbuilder">The custom AI agents: Elastic Agent Builder</h2>
<p>Beyond the native Elastic AI features, we built three bespoke AI agents using <a href="https://www.elastic.co/elasticsearch/agent-builder">Elastic Agent Builder</a>. Agent Builder is Elastic's framework for building custom AI agents that combine LLM instructions with modular, reusable tools, each tool being an ES|QL query, a built-in search capability, workflow execution, or an external integration via MCP. Agents parse natural language requests, select the appropriate tools, execute them, and iterate until they can provide a complete answer, all while managing context with data inside Elasticsearch. You can read more about the framework in the <a href="https://www.elastic.co/docs/explore-analyze/ai-features/elastic-agent-builder">Agent Builder documentation</a> and the <a href="https://www.elastic.co/search-labs/blog/elastic-ai-agent-builder-context-engineering-introduction">Elasticsearch Labs deep dive</a>.</p>
<p>The three key components of Agent Builder that we leveraged were:</p>
<p><strong>Agents:</strong> Custom LLM instructions and a set of assigned tools that define the agent's persona, capabilities, and behaviour boundaries. Each agent has a system prompt that controls its mission, the tools it can access, and the structure of its responses.</p>
<p><strong>Tools:</strong> Modular functions that agents use to search, retrieve, and manipulate Elasticsearch data. We built custom ES|QL tools that queried specific indices containing exercise documentation, playbooks, and reports.</p>
<p><strong>Agent Chat:</strong> The conversational interface - both the built-in Kibana UI and the programmatic API - that participants used to interact with the agents.</p>
<p>Agent and tool configurations are defined as JSON and managed via the Agent Builder APIs, making the entire agent lifecycle - from prompt engineering to tool binding - reproducible and version-controllable. We'll share the GrantPT agent configuration and tool definitions in a follow-up post for those who want to replicate this approach - watch this space.</p>
<p>Here's what each agent did:</p>
<h3 id="1grantptthegeneralpurposeassistant">1. GrantPT - The general-purpose assistant</h3>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt98943cb9194263ee/6a7d7f6dbd2198f7e175526b/image5.png" alt="" /></p>
<p>Available to all ~2,500 exercise participants, GrantPT was our primary AI agent and the best demonstration of how straightforward Agent Builder makes it to stand up a capable, domain-specific assistant. The agent's configuration consisted of a JSON object defining its system prompt, persona, and an array of bound tool IDs - that's it. No custom application code, no bespoke API layer, just declarative configuration.</p>
<p>What gave GrantPT its depth was the tooling. We defined a mix of built-in platform tools and custom ES|QL tools, each registered with a description, a parameterised query, and typed parameter definitions. For example, the knowledge base tool accepted a <code>target_index</code> and a semantic <code>query</code> parameter, executing a parameterised ES|QL query against our <code>dcm5-grantpt-*</code> indices with semantic search ranking:</p>
<pre><code>FROM dcm5-grantpt-* METADATA _score, _index
| WHERE _index == ?target_index
| WHERE content: ?query
| SORT _score DESC
| LIMIT 10
</code></pre>
<p>A separate index discovery tool let the agent dynamically enumerate available knowledge base indices at the start of each conversation, meaning we could add new documentation indices during the exercise without reconfiguring the agent; it would simply discover them on the next interaction.</p>
<p>We also built a Jira integration tool that performed semantic search across ingested helpdesk tickets, enabling GrantPT to surface relevant troubleshooting context from prior support requests. This was particularly useful for the HelpDesk Analysts, who could ask GrantPT about recurring issues and get responses grounded in actual ticket history rather than generic guidance.</p>
<p>The RBAC-tailored response behaviour came from a combination of the agent's system prompt, which instructed it to contextualise answers based on the user's role, and the underlying Elasticsearch security model. Because each tool's ES|QL query is executed within the user's security context, the agent can only surface documents accessible to the user's role. A Blue Team member asking about exercise procedures would get results scoped to their team's accessible indices, whilst a HelpDesk Analyst would see results from helpdesk-specific indices. The agent didn't need explicit role-switching logic; Elasticsearch's native document-level security handled scoping, and the agent simply worked with whatever results were returned. This is one of the things that makes Agent Builder genuinely elegant - by inheriting Elasticsearch's security model, you get RBAC-aware AI without writing a single line of authorisation code.</p>
<h3 id="2redrocktheadversaryscompanion">2. REDRock - The adversary's companion</h3>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt44c7c9b6fc1999ae/6a7d7f7042a117353b9590c6/image7.png" alt="" /></p>
<p>This agent was exclusively available to Red Teams. REDRock followed the same Agent Builder pattern, a dedicated system prompt defining its adversarial persona, bound to its own set of custom ES|QL tools querying Red Team-specific indices. These indices contained the Red Team playbooks, Tuoni C2 documentation, known system vulnerabilities within the range environment, and information about deployed services. The tool definitions mirrored the same parameterised semantic search pattern used by GrantPT, but were scoped to indices accessible only to Red Team roles. Red Team operators could query attack vectors, check for known weaknesses in target systems, and get contextual guidance on their operational plans. It was, quite frankly, like giving the attackers an extremely well-briefed operations officer.</p>
<h3 id="3refpttherefereestool">3. RefPT - The referee's tool</h3>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf04387af0aafed85/6a7d7f733cab1c29710e19b6/image2.png" alt="" /></p>
<p>Built specifically for the White Team (the exercise referees and assessors), RefPT was bound to tools querying indices containing Blue Team reports, scenario events, and the scoring criteria. Its purpose was to ensure uniform and fair scoring across all 40+ teams. The agent's system prompt was tuned to cross-reference submitted reports against known scenario events and scoring rubrics, helping assessors identify inconsistencies or gaps. When you've got assessors evaluating dozens of teams simultaneously, having an AI that can correlate reports against a structured scoring index is genuinely transformative for consistency.</p>
<h3 id="tinesaipoweredworkflowautomation">Tines: AI-powered workflow automation</h3>
<p>Tines was also a consumer of the on-range AI service. Each Blue Team had a dedicated Tines instance, with Tines action connectors provisioned in their Kibana space. Tines could leverage the Bedrock-backed AI capabilities for intelligent workflow automation, such as automated alert enrichment, AI-assisted triage decisions, natural-language summaries in notification workflows, and natural-language workflow creation. The Tines connector was configured per-team with credentials stored in Vault:</p>
<pre><code>resource "elasticstack_kibana_action_connector" "tines_bt" {
  count = var.team_count

  name              = "BT-${local.team_numbers[count.index]}-Tines"
  connector_type_id = ".tines"
  space_id          = local.space_ids[count.index]

  config = jsonencode({
    url = "https://tines.dsoc.${local.team_numbers[count.index]}.dcm.ex/"
  })
}
</code></pre>
<h3 id="ensuringcomplianceguardrailsandaudit">Ensuring compliance: Guardrails and audit</h3>
<p>Every AI interaction across all of these consumers was governed by strict AWS Bedrock Guardrails. We deployed guardrails with content filtering (hate, insults, sexual content, and violence at MEDIUM thresholds), PII protection (blocking email addresses, phone numbers, names, addresses, UK National Insurance numbers, credit card numbers, and IP addresses), topic-based filtering to prevent discussion of actual classified operations, and profanity filtering. Here's a snippet of the guardrail configuration from our Terraform:</p>
<pre><code>resource "aws_bedrock_guardrail" "dcm5_elastic" {
  name        = "dcm5-prod-elastic-guardrail"
  description = "Guardrails for DCM5 Prod Elastic Kibana GenAI connectors"

  content_policy_config {
    filters_config {
      input_strength  = "MEDIUM"
      output_strength = "MEDIUM"
      type            = "HATE"
    }
    # ... additional content filters for INSULTS, SEXUAL, VIOLENCE
  }

  sensitive_information_policy_config {
    pii_entities_config {
      action = "BLOCK"
      type   = "UK_NATIONAL_INSURANCE_NUMBER"
    }
    pii_entities_config {
      action = "BLOCK"
      type   = "IP_ADDRESS"
    }
    # ... additional PII filters
  }

  topic_policy_config {
    topics_config {
      name       = "classified-information"
      definition = "Discussions about actual classified operations, current real-world military activities, or operational intelligence."
      type       = "DENY"
    }
  }
}
</code></pre>
<p>Each Blue Team space had its own IAM user for Bedrock access, and the <code>genAiSettings:defaultAIConnectorOnly</code> Kibana setting was enforced to prevent teams from configuring their own connectors. This meant every single API call could be traced back to a specific team via CloudWatch, and the NSOC had complete audit visibility. The CloudWatch log group <code>/aws/bedrock/grantpt-prod/invocations</code> captured every invocation and guardrail event.</p>
<p>The numbers for all AI consumers speak for themselves: 3 custom AI Agents, 2,797 conversations, and 785 million AI tokens consumed throughout the exercise.</p>
<h2 id="ingamerealtimemonitoring">In-game real-time monitoring</h2>
<p>Within the exercise scenario, each team had access to RocketChat as their on-range messaging client. Every Blue Team got its own channel, the ability to direct message anyone in the exercise, and the freedom to spin up new channels as needed. Most critically for DCM tradition, this included the memes channel - the spiritual backbone of all inter-team ribbing and the creative morale-boosting humour that inevitably emerges when you put a few thousand cyber operators under pressure for a week.</p>
<p>All of this communication data represented a brilliant real-time window into range health, team sentiment, and the topics trending across the exercise. It felt too good to pass up, so we ingested the entire RocketChat conversation corpus into Elastic in real time and put it to work.</p>
<h3 id="sentimentanalysisandnamedentityrecognition">Sentiment analysis and named entity recognition</h3>
<p>For named entity recognition, we deployed the <a href="https://huggingface.co/dslim/bert-base-NER">dslim/bert-base-NER</a> model from Hugging Face into a machine learning node on the NSOC deployment using the <a href="https://www.elastic.co/guide/en/elasticsearch/client/eland/current/index.html">Elastic ELAND client</a>. This was then wired into an Elasticsearch ingest pipeline that every RocketChat message passed through on ingestion. We took the extracted entities and surfaced the most common ones as dashboard themes, giving us a live view of the ebb and flow of conversation topics throughout the exercise.</p>
<p>We also analysed group activity, user statistics, and general communication patterns to build a picture of life patterns for each team - most active participants, message volume over time, and sentiment trends pivoted by individual users. All told, it gave us some genuinely interesting insight into what was happening on the range in near real time. When we switched Elastic Agent into Prevent mode, for instance, a word cloud on our dashboard immediately lit up with "Elastic" as the most discussed theme across all channels - Blue Teams discussing its effectiveness, Red Teams lamenting their lost beacons. Rather satisfying, that.</p>
<h3 id="memeanalysisyesreally">Meme analysis (yes, really)</h3>
<p>Finally - and this one raised a few eyebrows - we pulled every meme submitted to the channels, vectorised the images, and ran nearest-neighbour evaluations to cluster similar memes and topics together. We also passed them through the zero-shot NER inference model to generate thematic descriptions of each meme's content. The logic was that these outputs might prove useful later for filtering, moderation, or other in-game interactions. Whether the meme analysis yielded operationally critical intelligence is debatable. Whether it was good fun is not.</p>
<h2 id="nippingproblemsinthebud">Nipping problems in the bud</h2>
<p>As much as we hoped everything would run smoothly during exercise week, things inevitably break, aren't fully understood, or need further customisation to suit how a particular team wants to use them. For this, we had our own subsection of the in-range helpdesk where Elastic and GenAI-specific requests could be raised by any team.</p>
<p>We manned this helpdesk for the entire duration of the exercise, providing guidance, documentation, issue debugging, and range-specific recommendations. That last point is worth expanding on. Sometimes, what a Blue Team was seeing in Elastic wasn't actually an Elastic problem at all, but rather Elastic faithfully surfacing something on the range that warranted further investigation (Red Teams can cause absolute mayhem, and the telemetry doesn't lie). Over the course of the exercise, we covered 125 individual support requests from teams specifically asking for help from us at Elastic.</p>
<h3 id="preemptivedebuggingwithtines">Pre-emptive debugging with Tines</h3>
<p>Beyond visiting teams via VTC or in person at EXCON, we also worked with <a href="https://www.tines.com/partners/elastic-security/">Tines</a> to try something a bit more proactive. We pulled the ticket body from incoming requests, attempted to categorise the problem, ran the categorisation against our corpus of previously resolved tickets, and had GenAI produce a summarised first-pass response aimed at solving the user's issue before triage brought it to our queue.</p>
<p>This is actually a pattern we borrowed from our own <a href="https://www.elastic.co/blog/elastic-wins-2025-best-use-of-ai-for-assisted-support">support organisation at Elastic</a>, where we provide a similar capability using our extensive knowledge base of previously solved issues as a repository for supporting AI Agent context. The idea is straightforward: use past solutions to give a machine-generated, informed first stab at resolving a problem, and short-circuit the need for a support engineer to pick up every ticket manually. It didn't solve everything; some issues genuinely needed a human with range context, but it meaningfully reduced the queue pressure and got faster answers to the teams who needed them. This was such a success with our own specific tickets and queue that we actually extended the remit to the entire helpdesk in the latter part of the exercise, helping to reduce the load on the other groups in the Green team supporting the exercise.</p>
<h2 id="industrypartnershipsbettertogether">Industry partnerships: Better together</h2>
<p>One of the things we're most proud of is how our partnership ecosystem has grown year on year. DCM is not just an Elastic show; it's a genuine coalition of industry partners, each bringing something unique to the security platform.</p>
<p><strong>Year 1 (DCM2)</strong> - Elastic joined as an industry partner, providing the security monitoring and endpoint detection platform.</p>
<p><strong>Year 2 (DCM3)</strong> - We brought in Endace, providing 1:1 packet capture capability. Full packet capture alongside Elastic's network visibility gave teams the ability to conduct deep-dive forensics that log-based analysis alone can't provide.</p>
<p><strong>Year 3 (DCM4)</strong> - Tines joined the family, bringing workflow automation to the table. Blue Teams could now build automated response playbooks, triage workflows, and notification chains, all integrated directly into their Elastic environment via the native Tines connector.</p>
<p><strong>Year 4 (DCM26, formerly DCM5)</strong> - AWS came on board, providing Bedrock access for our AI agents and contributing funding towards the Elastic deployments. This was a significant milestone; having a hyperscaler directly invested in the exercise's success unlocked capabilities (such as compliant, UK-tenanted AI inference with full guardrails and audit logging) that simply wouldn't have been possible otherwise. Tines' integration this year was also enhanced by the addition of on-range access to LLMs. The DCM series also reached a milestone this year, transitioning from its origins as an Army Cyber Association initiative to an officially funded programme under Cyber and Specialist Operations Command.</p>
<p><strong>To the teams at Endace, Tines, and AWS - sincere thanks. This exercise is better because of your contributions, and all Teams are better equipped because of the platform we've built together. We're already planning for DCM27. Cheers to the lot of you.</strong></p>
<h2 id="culturehighlightsandthebitsthatmakeitworthwhile">Culture, highlights, and the bits that make it worthwhile</h2>
<h3 id="thechallengecoins">The Challenge Coins</h3>
<p>We had custom challenge coins minted for DCM26. If you know, you know, challenge coins are a long-standing military tradition, and having one made for the exercise felt like the right way to mark our fourth year of involvement.</p>
<h3 id="thecocktailparty">The cocktail party</h3>
<p>We were also grateful to be invited to the High Commission cocktail party hosted by the British High Commissioner to Singapore. There's something quite surreal about discussing Elasticsearch shard counts and Terraform state management whilst holding a gin and tonic at the ambassador's invitation. It was a brilliant evening, a genuine reminder that these exercises exist at the intersection of technology and diplomacy, and that the relationships built here extend well beyond the technical.</p>
<h2 id="wrappingup">Wrapping up</h2>
<p>The multi-tenanted architecture proved itself under sustained load; the native Elastic AI features (<a href="https://www.elastic.co/elasticsearch/ai-assistant">AI Assistant</a> and <a href="https://www.elastic.co/docs/solutions/security/ai/attack-discovery">Attack Discovery</a>) gave teams capabilities that would have been science fiction a few years ago; and the custom AI agents exceeded our expectations for adoption. The partnership model continues to demonstrate that industry involvement in defence exercises creates outcomes that no single organisation could achieve alone.</p>
<p>Defence Cyber Marvel 2026 was a landmark iteration of an exercise that continues to grow in ambition, complexity, and impact. For Elastic, being trusted to provide the core defensive security platform for 40 Blue Teams from 29 nations, and this year, the AI capability as well, is something we don't take lightly. The exercise develops real skills for real people who will go on to defend real networks, and being a part of that mission is genuinely meaningful.</p>
<p>As the <a href="https://www.gov.uk/government/news/uk-to-lead-multinational-cyber-defence-exercise-from-singapore">UK Government's press release</a> put it, DCM demonstrates the practical value of real-life scenarios that reinforce international partnerships. We couldn't agree more.</p>
<p>We'll be back next year, and I suspect we'll have even more to talk about. In the meantime, we'll continue to improve the product so that support for environments such as Defence Cyber Marvel excels year over year.</p>
<p>See you on the range.</p>
<p>Follow the DCM26 story on social media:</p>
<p><a href="https://www.facebook.com/RSIGNALS/posts/last-week-defence-cyber-marvel-2026-based-in-singapore-brought-together-2500-par/1338105391677347/">Facebook</a> | <a href="https://www.linkedin.com/posts/uk-in-singapore_defence-cyber-marvel-2026pdf-activity-7426505462310752258-1aHq?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAABiQ31MBIbDwn5LYMrolM4rznGQcLabrY9A">LinkedIn</a> | <a href="https://www.instagram.com/p/DU00Y1jCKbr/">Instagram</a></p>
<h2 id="furtherreading">Further reading</h2>
<p><em>Elastic Security &amp; AI</em></p>
<ul>
<li><a href="https://www.elastic.co/security-labs">Elastic Security</a> - The platform powering the Blue Team deployments  </li>
<li><a href="https://www.elastic.co/elasticsearch/ai-assistant">AI Assistant for Security</a> - Context-aware AI chat within Elastic Security  </li>
<li><a href="https://www.elastic.co/docs/solutions/security/ai/attack-discovery">Attack Discovery</a> - LLM-powered alert correlation and threat narrative generation  </li>
<li><a href="https://www.elastic.co/docs/explore-analyze/ai-features/elastic-agent-builder">Agent Builder</a> - Framework for building custom AI agents with Elasticsearch</li>
</ul>
<p><em>Infrastructure &amp; Tooling</em></p>
<ul>
<li><a href="https://registry.terraform.io/providers/elastic/elasticstack/latest/docs">Elastic Stack Terraform Provider</a> - Infrastructure as Code for the Elastic Stack  </li>
<li><a href="https://www.elastic.co/docs/reference/fleet">Elastic Fleet Guide</a> - Centrally managing Elastic Agents at scale  </li>
<li><a href="https://github.com/ClarifiedSecurity/catapult">Catapult by Clarified Security</a> - Ansible-based cyber range provisioning</li>
</ul>
<p><em>Exercise Context</em></p>
<ul>
<li><a href="https://www.gov.uk/government/news/uk-to-lead-multinational-cyber-defence-exercise-from-singapore">UK Government DCM26 Press Release</a> - Official overview of the exercise</li>
</ul>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/elastic-defence-cyber-marvel</link>
    <guid isPermaLink="false">elastic-defence-cyber-marvel</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[James Garside]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb896872c76c9eb70/6a7d7f765967e57ae25da4f6/elastic-defence-cyber-marvel.webp" length="0" type="image/webp"/>
    <pubDate>Thu, 09 Apr 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[The Engineer's Guide to Elastic Detections as Code]]></title>
    <description><![CDATA[This post details the latest evolution of Elastic Security's Detections as Code (DaC) framework, including its development timeline, current feature highlights, and tailored implementation examples.]]></description>
    <content:encoded><![CDATA[<p>In an ever-evolving threat landscape, security operations are reaching a tipping point. As the velocity and complexity of threats increase, teams expand and managed environments multiply. Commonly, manual approaches to rule management become a bottleneck. This is where Detections as Code (DaC) steps in, not just as a tool, but as a methodology.</p>
<p>DaC as a methodology applies software development practices to the creation, management, and deployment of security detection rules. By treating detection rules as code, it enables version control, automated testing, and deployment processes, enhancing collaboration, consistency, and agility in response to threats. DaC streamlines the detection rule lifecycle, ensuring high-quality detections through peer reviews and automated tests. This methodology also supports compliance with change management requirements and fosters a mature security posture.</p>
<p>That's why we’re excited to share the latest updates to Elastic's <a href="https://github.com/elastic/detection-rules">detection-rules</a>, our open repository for writing, testing, and managing security detection rules in Elastic, that also allows you to create your own <a href="https://dac-reference.readthedocs.io/en/latest/">Detections as Code (DaC) framework</a>. Continue reading for highlighted implementation examples using extended functionality, and the announcement of Elastic's free Detections as Code Workshop.</p>
<h2 id="elasticsecuritydacthejourneyfromalphatogeneralavailability">Elastic Security DaC: The journey from alpha to general availability</h2>
<p>With the functionality now provided in <a href="https://github.com/elastic/detection-rules">detection-rules</a> repository, users can manage all their detection rules as code, review rule tunings, automatically test and validate rules, and automate rules deployment across their environments.</p>
<h3 id="pre2024elasticsinternaluseofdac">Pre-2024: Elastic’s internal use of DaC</h3>
<p>Elastic threat research and detection engineering team created and used the <a href="https://github.com/elastic/detection-rules">detection-rules</a> repository to develop, test, manage and release prebuilt rules, following DaC principles - reviewing rules as a team, automating their tests and release. The repository also has an interactive CLI to create rules, so engineers could start working on the rules right there. </p>
<p>As the security community's interest in as-code principles grew, and the available Elastic Security APIs already allowed users to implement their custom Detections as code solutions, Elastic decided to extend the <a href="https://github.com/elastic/detection-rules">detection-rules</a> repository functionality to enable our users to benefit from our tooling and aid them in creating their DaC processes. </p>
<p>Here are the key milestones of Elastic’s user-focused DaC development from alpha to general availability.</p>
<h3 id="may2024alphareleaseofnewrollyourownfeatures">May 2024: Alpha release of new "roll your own” features</h3>
<p>Our detection-rules repository is adjusted for customer use, allowing for managing custom rules, adapting the test suite for user needs, and allowing for management of actions and exceptions alongside the rules.</p>
<p>Key additions:</p>
<ul>
<li>Custom rules directory support  </li>
<li>Select which test to run based on your requirements  </li>
<li><a href="https://www.elastic.co/docs/solutions/security/detect-and-alert/rule-exceptions">Exceptions</a> and Actions support</li>
</ul>
<p>We also published an extensive <a href="https://dac-reference.readthedocs.io/en/latest/">guidance</a> for Detections as Code with examples of implementation with Elastic Security using <a href="https://github.com/elastic/detection-rules">detection-rules</a> repository.</p>
<h3 id="august2024rollyourownfeaturesnowbeta">August 2024: "Roll your own” features now beta</h3>
<p>The functionality is extended to allow import and export of custom rules between Elastic Security and repository, more configuration options and versioning functionality extended to custom rules.</p>
<p>New features added:</p>
<ul>
<li>Bulk import/export of custom rules (based on Elastic Security APIs)   </li>
<li>Fully configurable unit test, validation, and schemas  </li>
<li>Version lock for custom rules</li>
</ul>
<h3 id="marchaugust2025aregenerallyavailableandsupported">March - August 2025: are generally available and supported</h3>
<p>Using DaC with Elastic Security 8.18 and up:</p>
<ul>
<li><a href="https://www.elastic.co/guide/en/security/8.18/whats-new.html#_customize_and_manage_prebuilt_detection_rules">Supports prebuilt rules management</a>. You can export all prebuilt rules from Elastic Security and store them alongside your custom rules.  </li>
<li>Support for rules filtering for export added.</li>
</ul>
<p>Adjacent to DaC efforts, we also released new Terraform resources (<a href="https://github.com/elastic/terraform-provider-elasticstack/releases/tag/v0.12.0">V0.12.0</a>  and <a href="https://github.com/elastic/terraform-provider-elasticstack/releases/tag/v0.13.0">V0.13.0</a>) in October-December 2025, allowing Terraform users to manage detection rules and exceptions.</p>
<p>With this foundation spelled out, let's explore the powerful features that are available to streamline your detection engineering process.</p>
<h2 id="detectionrulesdacfunctionalityhighlights">Detection-rules DaC functionality highlights</h2>
<p>There are a few worthwhile additions since our <a href="https://www.elastic.co/security-labs/dac-beta-release">last DaC publication</a>, which we’ll expand on below.</p>
<h3 id="additionalfilters">Additional filters</h3>
<p>The <a href="https://github.com/elastic/detection-rules/blob/main/CLI.md#exporting-rules">filter functionality</a> available when exporting rules from Kibana has been extended to allow you to precisely define which rules to sync in DaC. Here are the new flags:</p>
<p>| Flag | Description |
| :---: | ----- |
| <strong>-cro</strong> | Filters the export to only include rules created by the user (not Elastic prebuilt rules). |
| <strong>-eq</strong> | Applies a query filter to the rules being exported. |</p>
<p>Let’s take an example of when you wish to organize rules by data source, and want to export the AWS rules to a specific folder. In this case, let’s use filtering on tags for data sources and export all rules with the <code>Data Source AWS</code> tag:</p>
<pre><code>python -m detection_rules kibana export-rules -d dac_test/rules #add rules to the dac_test/rules folder
-sv #strip the version fields from all rules
-cro #export only custom rules
-eq "alert.attributes.tags: "Data Source: AWS"" # export only rules with "Data Source: AWS" tag
</code></pre>
<p>See Kibana documentation for <a href="https://www.elastic.co/docs/api/doc/kibana/operation/operation-performrulesbulkaction#operation-performrulesbulkaction-body-application-json-query">query string filtering</a> for the underlying API call used here and the <a href="https://www.elastic.co/docs/api/doc/kibana/operation/operation-findrules">list all detection rules API call</a> for example available fields to construct the query filter.</p>
<h3 id="customfolderstructure">Custom folder structure</h3>
<p>In the detection-rules repo, we use a folder structure based on platform, integration, and MITRE ATT&amp;CK information. This helps us with our organization and rule development. This is by no means the only method of organization. You may want to organize your rules by customer, date, or source as examples. This will vary greatly depending on your use case.</p>
<p>Whether you use this export process or manual organization, once you have your rules in a location or folder structure that you like, you can now keep this local structure even when re-exporting rules. It is important to note that the new rules need to be placed in their desired location manually. The local rule-loading mechanism detects where the rules are placed in order to know where to put them. If the rule is not there, it will then use the specified output directory to place the new rule(s). To use the local rule loading for updating existing rules use the <code>--load-rule-loading / -lr</code> flag for the <code>kibana export-rules</code> and <code>import-rules-to-repo</code> commands. These flags enable you to make use of the local folders specified in your <code>config.yaml</code>.</p>
<p>Let’s look at example with the rules organised in folders the following way:</p>
<p><code>rule_dirs:</code><br />
  <code>- rules</code><br />
    <code>my_test_rule.toml</code><br />
  <code>- another_rules_dir</code><br />
    <code>high_number_of_process_and_or_service_terminations.toml</code></p>
<p>We’ll specify the following in the <code>config.yaml</code> file:</p>
<p><code>rule_dirs:</code><br />
  <code>- rules</code><br />
  <code>- another_rules_dir</code></p>
<p>With the new <code>-lr</code> option, rule updates from Kibana will now use these additional paths instead of exporting directly to the specified directory.</p>
<p>Running <code>python -m detection_rules kibana --space test_local export-rules -d dac_test/rules/ -sv -ac -e -lr,</code>will export rules from <code>test_local</code>  space,  <code>my_test_rule.toml</code> will be written to dac_test/rules/ as it was already on disk there and <code>high_number_of_process_and_or_service_terminations.toml</code> will be written to <code>dac_test/another_rules_dir/.</code></p>
<p>This can be particularly useful if you have the same rules in different sub-folder configurations for different customers. For example, let’s say you have your rules broken down by platform and integration similar to Elastic’s prebuilt rule folder structure. For your customers, SOCs, or threat-hunting teams, having the rules organized underneath these platform/integration folders may be the most useful mechanism for them to manage the rules. However, your information security team or primary detection engineering team may want to manage the rules by initiative or rule author instead so that all the rules a particular individual or team is responsible for are organized in one place. Now with the local rule-loading flags, you can simply have two configuration files and the duplicated rules in each structure. When you are exporting updates for the rules, you would then use the environment variable to select the appropriate configuration file and export the rule updates. These updates will then be applied to the rules in place, maintaining the directory structure.   </p>
<h3 id="miscellaneouslocalloadingupdates">Miscellaneous local loading updates</h3>
<p>In addition to the above, we have added two smaller new features designed to help users who are adding local information in the detection rules TOML files and schema. These are as follows: </p>
<ol>
<li>Local date support from the local files where the local date will be maintained from the original file   </li>
<li>Upgrades to the auto gen feature to inherit known types from existing schema.</li>
</ol>
<p>The local date component can be useful when one wants more manual control over the date field in the file. Without using the override, the date will be based on when the Kibana rule contents were exported. Using the <code>--local-creation-date</code> flag, the date will not be updated when the file contents are re-exported. </p>
<p>The automatic schema generation has been updated to inherit the types from other indices/integrations if they are present. This provides a potentially more accurate schema, as well as reducing the need for manual updates after the fact. For example, you have a rule that uses the index “new-integration*” with the following fields:</p>
<ul>
<li><code>host.os.type.new_field</code>  </li>
<li><code>dll.Ext.relative_file_creation_time</code>  </li>
<li><code>process.name.okta.thread</code></li>
</ul>
<p>Instead of each of these fields being added to the schema with a default type, their types are inherited from existing schemas. In this case, the types for <code>dll.Ext.relative_file_creation_time</code> and <code>process.name.okta.thread</code> are inherited. </p>
<pre><code>{
  "new-integration*": {
    "dll.Ext.relative_file_creation_time": "double",
    "host.os.type.new_field": "keyword",
    "process.name.okta.thread": "keyword"
  }
}
</code></pre>
<p>To see how to use this with your custom data types, see the <a href="https://www.elastic.co/security-labs/blog/detection-as-code-timeline-and-new-features#Custom-schemas-usage">Custom schemas usage</a> section within the Implementation examples part of this blog. </p>
<h2 id="expandingonusageexamples">Expanding on usage examples</h2>
<p>Below you will find more examples of DaC implementations, these are not focused on new functionality additions, but go deeper on the topics we see discussed in the community.</p>
<p>It’s worth noting that Detections as Code features are provided as components that can be used to build a custom implementation for your chosen process and architecture. When implementing DaC in your production environment, treat it as an engineering process and follow <a href="https://dac-reference.readthedocs.io/en/latest/dac_concept_and_workflows.html#best-practices">the best practices.</a></p>
<h3 id="dacimplementationwithgitlab">DaC implementation with Gitlab</h3>
<p>When we look at implementations of DaC typically this revolves around using some form of CI/CD product to automatically perform rule management based on a given trigger. These triggers vary considerably based on the desired setup, specifically the authoritative source of rules and the desired state of your version control system (VCS). For a much more in-depth exploration of some of these considerations, see our <a href="https://dac-reference.readthedocs.io/en/latest/core_component_syncing_rules_and_data_from_vcs_to_elastic_security.html">DaC Reference Material</a>. Below is a simple example using Gitlab as VCS provider and using its in-built CI/CD via Gitlab Actions.     </p>
<pre><code>stages:                # Define the pipeline stages
  - sync               # Add a 'sync' stage

sync-to-production:    # Define a job named 'sync-to-production'
  stage: sync          # Assign this job to the 'sync' stage
  image: python:3.12   # Use the Python 3.12 Docker image
  variables:
    CUSTOM_RULES_DIR: $CUSTOM_RULES_DIR    # Set custom rules env var
  script:                                  # List of commands to run 
    - python -m pip install --upgrade pip  # Upgrade pip
    - pip cache purge                      # Clear pip cache
    - pip install .[dev]                   # Install package w/ dev deps
    - |  # Multi-line command to import rules                                        
      FLAGS="-d ${CUSTOM_RULES_DIR}/rules/ --overwrite -e -ac"
      python -m detection_rules kibana --space production import-rules $FLAGS
  environment:
    name: production   # Specify deployment environment as 'production'
  only:
    refs:
      - main           # Run this job only on the 'main' branch
    changes:
      - '**/*.toml'    # Run this job only if .toml files have changed
</code></pre>
<p>This is very similar to other inbuilt CI/CD from other Git-based VCS like Gitlab and Gitea. The main difference being in the syntax determining the triggering event. The DaC commands such as <code>kibana import-rules</code> would be the same regardless of VCS. In this example, we are syncing rules from our fork of the detection-rules repo to our Kibana Production Space. This is based on a number of prior decisions being made, for instance requiring unit tests to pass before merging rule updates and that rules on main being ready for prod. For a Github-based walkthrough of these considerations for this particular approach, please take a look at our <a href="https://dac-reference.readthedocs.io/en/latest/etoe_reference_example.html#demo-video">demo video</a>. </p>
<h3 id="customunittestingtipsandexamples">Custom Unit Testing tips and examples</h3>
<p>When considering DaC as a capability to add to your detection toolkit, setting up the CI/CD and base infrastructure should be considered as the first step in an ongoing process to improve the quality and usefulness of your rules. One of the key purposes in having “as code” tooling is adding the ability to further customize tooling to your needs and environment.  </p>
<p>One example of this is unit testing for rules. Beyond base functionality testing, some other key existing unit tests enforce Elastic-specific considerations around rule performance and optimization, as well as organization of metadata and tagging. This helps detection engineers and threat researchers remain consistent in their rule development. Building on this example, one may want to consider adding custom unit tests based on your specific needs.  </p>
<p>To illustrate this, take a Security Operations Center (SOC) environment where there are a number of analysts responsible for various different domains and tasks. When an alert is raised in the SIEM, it may not be immediately obvious who should handle remediation, or what team(s) need to be informed of the incident. Tagging the rules with a team tag: e.g. <code>Team: Windows Servers</code> similarly to how Elastic uses tags for data sources, can provide the SOC with a point of contact directly in the alert for who can help with remediation. </p>
<p>In our DaC environment, we can quickly create a new testing module to enforce this on all of the custom rules (or pre-built too). For this test, we are going to enforce having a <code>Team: &lt;some name&gt;</code> tag on all production rules that are not authored by Elastic. In the detection-rules repo, our testing is handled through the Python test suite called <code>pytest</code> and as such unit tests are organized into python modules (files) and subsequent classes and functions in these files under the <code>tests/</code> folder. To add tests simply either add classes or functions to the existing files or create a new one. In general, we recommend creating new test files so that you can receive updates to the existing tests from Elastic without having to merge the differences.</p>
<p>We will start by creating a new python file called <code>test_custom_rules.py</code> in the <code>tests/</code> directory with the following contents:</p>
<pre><code># test_custom_rules.py

"""Unit Tests for Custom Rules."""

from .base import BaseRuleTest


class TestCustomRules(BaseRuleTest):
    """Test custom rules for given criteria."""

    def test_custom_rule_team_tag(self):
        """Unit test that all custom rules have a Team: &lt;team_name&gt; tag."""
        tag_format = "Team: &lt;team_name&gt;"
        for rule in self.all_rules:
            if "Elastic" not in rule.contents.data.author:
                tags = rule.contents.data.tags
                if tags:
                    self.assertTrue(
                        any(tag.startswith("Team: ") for tag in tags),
                        f"Custom rule {rule.contents.data.rule_id} does not have a {tag_format} tag",
                    )
                else:
                    raise AssertionError(
                        f"Custom rule {rule.contents.data.rule_id} does not have any tags, include a {tag_format} tag"
                    )
</code></pre>
<p>Now each non-Elastic rule will be required to have a tag in the specified pattern for a team responsible for remediation. E.g. <code>Team: Team A.</code></p>
<h3 id="customschemasusage">Custom schemas usage</h3>
<p>Elastic’s ability to bring your own data types also extends to our DaC capabilities. For example, let’s take a look at some custom schemas for network protocols. Diverse data you have in your stack can of course be queried by your rules, and we will also want to leverage the applicable validation and testing for any custom rules on these data types too. This is where Custom schemas come in handy.</p>
<p>When we are validating queries, the query is parsed into the respective fields and the types of these fields are compared against what is provided in a given schema (e.g. <a href="https://www.elastic.co/docs/reference/ecs/ecs-field-reference">ECS schema</a>, the AWS Integration for AWS data, etc.). For custom data types, this follows the same validation path, with the ability to pull from locally defined custom schemas. These schema files can be built by hand as one or more json files; however, if you have some sample data already in your stack, you can take advantage of this and use it as validation and generate your schemas automatically. </p>
<p>Assuming you already have a custom rules folder configured (if not see instructions), you can turn on automatic schema generation by adding <code>auto_gen_schema_file: &lt;path_to_your_json_file&gt;</code> to your config file. This will generate a schema file in the specified location that will be used to add entries for each field and index combination. The file will be updated during any command where rule contents are validated against a schema, including import-rules-to-repo, kibana export-rules, view-rule, and others. This will also automatically add it to your stack-schema-map.yaml file when using a custom rules directory and config.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf66f76c382b77fad/6a7d7f27227b1c592c5957ac/image1.gif" alt="" /></p>
<p>With this power comes an increased responsibility on rule reviewers as any field used in the query is immediately assumed to be valid and added to the schema. One way to mitigate risk is to utilize a development space that has access to the data. In the PR, one can then link to a successful execution of the query with stack level validation on its data types. Once this is approved, one can remove the <code>auto_gen_schema_file</code> addition to the config and you now have a known valid schema based on your custom data. This provides a baseline for other rule authors to build upon as needed and maintains the type checking validation.</p>
<h2 id="learnmoreaboutdacandtryityourself">Learn more about DaC and try it yourself</h2>
<p>You can experience Elastic Security's Detections as Code (DaC) functionality firsthand with our interactive <a href="https://play.instruqt.com/elastic/invite/uqlknuayvxhy">Instruqt training</a>. This training provides a straightforward way to explore core DaC features in a pre-configured test environment, eliminating the need for manual setup. Give it a try!</p>
<p>If you are implementing DaC, share your experience, ask your questions and help others on the community slack <a href="https://elasticstack.slack.com/archives/C06TE19EP09">DaC channel</a>.</p>
<h2 id="trialelasticsecurity">Trial Elastic Security</h2>
<p>To experience the full benefits of what Elastic has to offer for detection engineers, start your Elastic Security <a href="https://cloud.elastic.co/registration">free trial</a>. Visit <a href="https://www.elastic.co/security">elastic.co/security</a> to learn more.</p>
<p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/detection-as-code-timeline-and-new-features</link>
    <guid isPermaLink="false">detection-as-code-timeline-and-new-features</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Eric Forte,Kseniia Ignatovych]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt91be65ac9e4c46a2/6a7d7f2b3ce8e27525cf2612/image2.png" length="0" type="image/png"/>
    <pubDate>Wed, 04 Feb 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Time-to-Patch Metrics: A Survival Analysis Approach Using Qualys and Elastic]]></title>
    <description><![CDATA[In this article, we describe how we applied survival analysis to vulnerability management (VM) data from Qualys VMDR, using the Elastic Stack.]]></description>
    <content:encoded><![CDATA[<p>Understanding how quickly vulnerabilities are remediated across different environments and teams is critical to maintaining a strong security posture. In this article, we describe how we applied <strong>survival analysis</strong> to vulnerability management (VM) data from <strong>Qualys VMDR</strong>, using the <strong>Elastic Stack</strong>. This allowed us to not only confirm general assumptions about team velocity (how quickly teams complete work) and remediation capacity (how much fixing they can take on) but also derive measurable insights. Since most of our security data is in the Elastic Stack, this process should be easily reproducible to other security data sources. </p>
<h2 id="whywedidit">Why We Did It</h2>
<p>Our primary motivation was to <strong>move from general assumptions to data-backed insights</strong> about:</p>
<ul>
<li>How quickly different teams and environments patch vulnerabilities  </li>
<li>Whether patching performance meets internal service level objectives (SLOs)  </li>
<li>Where bottlenecks or delays commonly occur  </li>
<li>What other factors can affect patching performance</li>
</ul>
<h2 id="whysurvivalanalysisabetteralternativetomeantimetoremediate">Why Survival Analysis? A Better Alternative to Mean Time to Remediate</h2>
<p>Mean Time to Remediate (MTTR) is commonly used to track how quickly vulnerabilities are patched, but both the mean and median suffer from significant limitations (we provide an example later in this article). The mean is highly sensitive to <em>outliers</em>[^1] and assumes the remediation times are evenly balanced around the average remediation time, which is rarely the case in practice. The median is less sensitive to extremes but discards information about the shape of the distribution and says nothing about the long tail of slow-to-patch vulnerabilities. Neither accounts for unresolved cases, i.e. vulnerabilities that remain open beyond the observation window, which are often excluded entirely. In practice, the vulnerabilities that remain open the longest are precisely the ones we should be most concerned about.</p>
<p><strong>Survival analysis</strong> addresses these limitations. Originating in medical and actuarial contexts, it models <strong>time-to-event data</strong> while explicitly incorporating <strong>censored observations</strong>, meaning in our context vulnerabilities that remain open. (For more details on its application to vulnerability management we strongly recommend <a href="https://www.themetricsmanifesto.com">“The Metrics Manifesto”</a>). Instead of collapsing remediation behavior into a single number, survival analysis estimates the probability that a vulnerability remains unpatched over time (e.g. 90% of vulnerabilities are remediated within 30 days). This allows for more meaningful assessments, such as the proportion of vulnerabilities patched within SLO (for example within 30, 90, or 180 days).</p>
<p>Survival analysis provides us with a <strong>survival function</strong> that estimates the probability a vulnerability remains unpatched over time.</p>
<p>::: 
This method offers a better view of remediation performance, allowing us to assess not just how long vulnerabilities persist, but also how remediation behavior differs across systems, teams, or severity levels. It’s particularly well-suited to security data, which is often incomplete, skewed, and resistant to assumptions of normality.
:::</p>
<h2 id="context">Context</h2>
<p>Although we have applied survival analysis across different environments, teams and organizations, in this blog we focus on the results for the Elastic Cloud production environment. </p>
<h3 id="vulnerabilityagecalculation">Vulnerability age calculation</h3>
<p>There are different methods to calculate vulnerability age.</p>
<p>For our internal metrics like <a href="https://www.elastic.co/blog/how-infosec-uses-elastic-stack-vulnerability-management">vulnerability adherence SLO</a>, we define vulnerability age as the difference between when a vulnerability was last found and when it was first detected (usually a few days after publication). This approach aims to penalize vulnerabilities that are reintroduced from an outdated base image. In the past, our base images were not updated frequently enough for our satisfaction. If a new instance is created, vulnerabilities can have a significant age (e.g., 100 days) from day one of discovery.</p>
<p>For this analysis, we find it more relevant to calculate the age based on the number of days between the last found date and the first found date. In this case, age represents the number of days the system was effectively exposed.</p>
<h3 id="patcheverythingstrategy">“Patch everything” strategy</h3>
<p>In our Cloud environment, we maintain a policy to patch everything. This is because we almost exclusively use the same base image across all instances. Since Elastic Cloud operates fully on containers, there are no specific application packages (e.g., Elasticsearch) installed directly on our systems. Our fleet remains homogeneous as a result. </p>
<h2 id="datapipeline">Data Pipeline</h2>
<p>Ingesting and mapping data into the Elastic Stack can be cumbersome. Luckily, we have <a href="https://www.elastic.co/integrations/data-integrations?solution=all-solutions&amp;category=security">many security integrations</a> that handle those natively, <a href="https://www.elastic.co/docs/reference/integrations/qualys_vmdr">Qualys VMDR</a> being one of them.  </p>
<p>This integration has 3 main interests over custom ingestion methods (e.g. scripts, beats, …):</p>
<ul>
<li>It natively enriches vulnerability data from the Qualys Knowledge Base which add CVE IDs, threat intel information, … <strong>without needing to configure enrich pipelines</strong>.  </li>
<li>Qualys data is already mapped to the Elastic Common Schema which is a standardized way of representing data, whether it’s coming from one source or another: for example, CVEs are always stored in field <a href="http://vulnerability.id"><em>vulnerability.id</em></a>, independent of the source.   </li>
<li>A transform with the latest vulnerability is already set up. This index can be queried to get the latest vulnerabilities status. </li>
</ul>
<h3 id="qualysagentintegrationconfiguration">Qualys agent integration configuration</h3>
<p>For survival analysis, we need to ingest both active and patched vulnerabilities. To analyze a specific period, we need to set the number of days in field <code>max_days_since_detection_updated</code>. In our environment, we ingest Qualys data daily, so there’s no need to ingest a long history of fixed data, as we’ve already done that.  </p>
<p>The Qualys VMDR elastic agent integration has been configured with the following:</p>
<p>| Property | Value | Comment |
| :---- | :---- | :---- |
| (Settings section) Username |  |  |
| (Settings section) Password |  | Since there are no API keys available in Qualys, we can only authenticate with Basic Authentication.  Make sure SSO is disabled on this account |
| URL | <a href="https://qualysapi.qg2.apps.qualys.com">https://qualysapi.qg2.apps.qualys.com</a> (for US2) | <a href="https://www.qualys.com/platform-identification/">https://www.qualys.com/platform-identification/</a>  |
| Interval | 4h | Adjust it based on the number of ingested events.  |
| Input parameters | show_asset_id=1&amp; include_vuln_type=confirmed&amp;show_results=1&amp;max_days_since_detection_updated=3&amp;status=New,Active,Re-Opened,Fixed&amp;filter_superseded_qids=1&amp;use_tags=1&amp;tag_set_by=name&amp;tag_include_selector=all&amp;tag_exclude_selector=any&amp;tag_set_include=status:running&amp;tag_set_exclude=status:terminated,status:stopped,status:stale&amp;show_tags=1&amp;show_cloud_tags=1 | show_asset_id=1: retrieve asset id show_results=1: details about what is the current installed package and which version should be installed max_days_since_detection_updated=3: filter out any vulnerabilities that haven’t been updated over the last 3 days (e.g. patched older than 3 days) status=New,Active,Re-Opened,Fixed: all vulnerability status are ingested filter_superseded_qids=1: ignore superseded ‘vulnerabilities Tags: filter by tags show_tags=1: retrieve Qualys tags show_cloud_tags=1: retrieve Cloud tags |</p>
<p>Once data is fully ingested, it can be reviewed either in Kibana Discover (logs-* data view -&gt; <em>data_stream.dataset : "qualys_vmdr.asset_host_detection"</em> ), either in the Kibana Security App (Findings -&gt; Vulnerabilities).</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd916b27bf56ffc40/6a7d8600fc63ab634564a044/image6.png" alt="" /></p>
<h3 id="loadingdataintopythonwiththeelasticsearchclient">Loading data into Python with the elasticsearch client</h3>
<p>Since the survival analysis calculation will be done in Python, we need to extract data from elastic into a python dataframe. There are several ways to achieve this, and in this article we’ll focus on two of them.</p>
<h4 id="withesql">With ES|QL</h4>
<p>The easiest and most convenient way is to leverage ES|QL with the arrow format. It’ll automatically populate the python dataframe (rows and columns). We recommend reading the blog post <a href="https://www.elastic.co/search-labs/blog/esql-pandas-native-dataframes-python">From ES|QL to native Pandas dataframes in Python</a> to get more details.  </p>
<pre><code>from elasticsearch import Elasticsearch
import pandas as pd

client = Elasticsearch(
    "https://[host].elastic-cloud.com",
    api_key="...",
)

response = client.esql.query(
    query="""
   FROM logs-qualys_vmdr.asset_host_detection-default
    | WHERE elastic.owner.team == "platform-security" AND elastic.environment == "production"
    | WHERE qualys_vmdr.asset_host_detection.vulnerability.is_ignored == FALSE
    | EVAL vulnerability_age = DATE_DIFF("day", qualys_vmdr.asset_host_detection.vulnerability.first_found_datetime, qualys_vmdr.asset_host_detection.vulnerability.last_found_datetime)
    | STATS 
        mean=AVG(vulnerability_age), 
        median=MEDIAN(vulnerability_age)
    """,
    format="arrow",
)
df = response.to_pandas(types_mapper=pd.ArrowDtype)
print(df)
</code></pre>
<p>Today, we have a limitation with ESQL: we can’t paginate through results. Therefore we are limited to 10K output documents (100K if server configuration is modified). Progress can be followed through this <a href="https://github.com/elastic/elasticsearch/issues/100000">enhancement request</a>. </p>
<h4 id="withdsl">With DSL</h4>
<p>In the elasticsearch python client, there is a native feature to extract all the data from a query with transparent pagination. The challenging part is to create the DSL query. We recommend creating the query in Discover and then click on Inspect, and then Request tab to get the DSL query. </p>
<pre><code>query = {
    "track_total_hits": True,
    "query": {
        "bool": {
            "filter": [
                {
                    "match": {
                        "elastic.owner.team": "awesome-sre-team"
                    }
                },
                {
                    "match": {
                        "elastic.environment": "production"
                    }
                },
                {
                    "match": {
"qualys_vmdr.asset_host_detection.vulnerability.is_ignored": False
                    }
                }
            ]
        }
    },
    "fields": [
        "@timestamp",
        "qualys_vmdr.asset_host_detection.vulnerability.unique_vuln_id",
        "qualys_vmdr.asset_host_detection.vulnerability.first_found_datetime",
        "qualys_vmdr.asset_host_detection.vulnerability.last_found_datetime",
        "elastic.vulnerability.age",
        "qualys_vmdr.asset_host_detection.vulnerability.status",
        "vulnerability.severity",
        "qualys_vmdr.asset_host_detection.vulnerability.is_ignored"
    ],
    "_source": False
}

results = list(scan(
        client=es,
        query=query,
        scroll='30m',
        index=source_index,
        size=10000,
        raise_on_error=True,
        preserve_order=False,
        clear_scroll=True
    ))
</code></pre>
<h2 id="survivalanalysis">Survival Analysis</h2>
<p>You can refer to the <a href="https://github.com/lauravoicu/elastic-vm-survivalanalysis/tree/main">code</a> to understand or reproduce it on your dataset. </p>
<h2 id="whatwelearned">What We Learned</h2>
<p>Leaning in on the research from the <a href="https://www.cyentia.com/why-your-mttr-is-probably-bogus/">Cyentia Institute</a> we looked at a few different ways to measure how long it takes to remediate vulnerabilities using means, medians, and survival curves. Each method gives a different lens through which we can understand time-to-patch data, and the comparison is important because depending on which method we use, we would draw very different conclusions about how well vulnerabilities are being addressed.</p>
<p>The first method focuses only on vulnerabilities that have already been closed. It calculates the median and mean time it took to patch them. This is intuitive and simple, but it leaves out a potentially large and important portion of the data (the vulnerabilities that are still open). As a result, it tends to underestimate the true time it takes to remediate, especially if some vulnerabilities stay open much longer than others.</p>
<p>The second method tries to include both closed and open vulnerabilities by using the time they’ve been open <em>so far</em>. There are many options to approximate a time-to-patch for the open vulnerabilities, but for simplicity here we assumed they were (will be?) patched at the time of reporting, which we know isn’t true. But it does offer a way to factor in their existence.</p>
<p>The third method uses survival analysis. Specifically, we used the Kaplan-Meier estimator to model the likelihood that a vulnerability is still open at any given time. This method handles the open vulnerabilities properly: instead of pretending they’re patched, it treats them as “censored” data. The survival curve it produces drops over time, showing the proportion of vulnerabilities still open as days or weeks pass. </p>
<h3 id="howlongdovulnerabilitieslast">How Long Do Vulnerabilities Last?</h3>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4ae0aedea070cc7e/6a7d860373d9bd27d729acaa/image4.png" alt="" /></p>
<p>In the current 6-month snapshot[^2], the closed-only time-to-patch has a median ~33 days and a mean ~35 days. On the surface that looks reasonable, but the Kaplan-Meier curve shows what those numbers hide: at 33 days, ~54% are still open; at 35 days, ~46% are still open. So even around the “typical” one-month mark, about half of issues remain unresolved.</p>
<p>We also computed observed-so-far statistics (treating open vulnerabilities as if they were patched at the end of the measurement window). In this window they happen to be almost the same (median ~33 days, mean ~35 days) because the ages of today’s open items cluster near one month. That coincidence can make averages look reassuring, but it’s incidental and unstable: if we shift the snapshot to just before the monthly patch push and these same statistics drop sharply (we’ve seen an observed median of ~19 days and observed a mean of ~15 days) without any change in the underlying process.</p>
<p>The survival curve avoids that trap, because it answers the question of “% still open after 30/60/90 days”, and offers visibility into the long tail that stays open well past a month.</p>
<h3 id="patcheverythingeverywherethesameway">Patch Everything Everywhere The Same Way?</h3>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2d33872ebab47e81/6a7d860663e9594b0173af31/image5.png" alt="" /></p>
<p>Stratified survival analysis takes the idea of survival curves one step further. Instead of looking at all vulnerabilities together in one big pool, it separates them into groups (or “strata”) based on some meaningful characteristic. In our analysis, we have stratified vulnerabilities by severity, asset criticality, environment, cloud provider, team/division/organization. Each group gets its own survival curve, and here in the example graph we compare how quickly different vulnerability severities are remediated over time.</p>
<p>The benefit of this approach is that it exposes differences that would otherwise be hidden in the aggregate. If we only looked at the overall survival curve, we can only make conclusions about the remediation performance across the board. But stratification reveals if different teams, environments or severity issues are addressed faster than the rest, and in our case that the patch everything strategy is indeed consistent. This level of detail is important for making targeted improvements, helping us understand not just how long remediation takes in general, but if and where real bottlenecks exist.</p>
<h3 id="howfastdoteamsact">How Fast Do Teams Act?</h3>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta52fdeb6cc2cefb5/6a7d860ae3a2192f1099c7a4/image2.png" alt="" /></p>
<p>While the survival curve emphasizes how long vulnerabilities remain open, we can flip the perspective by using the cumulative distribution function (CDF) instead. The CDF focuses on how quickly vulnerabilities are patched, showing the proportion of vulnerabilities that have been remediated by a given point in time.</p>
<p>Our choice of plotting the CDF provides a clear picture of remediation speed, however it’s important to note that this version includes only vulnerabilities that were patched within the observed time window. Unlike the survival curve which we compute over a rolling 6-month cohort to capture full lifecycles, the CDF is computed month-over-month on items closed in that month[^3]. </p>
<p>As such, it tells us how quickly teams remediate vulnerabilities <strong>once they do so</strong>, and it doesn’t reflect how long unresolved vulnerabilities remain open. For example, we see that 83.2% of the vulnerabilities closed in the current month were resolved within 30 days of the first detection. This highlights patching velocity for recent, successful patches but does not account for longer-standing vulnerabilities that remain open and are likely to have longer time-to-patch durations. Therefore, we use the CDF for understanding short-term response behavior, whereas the full lifecycle dynamics are given by a combination of CDF alongside survival analysis: the CDF describes <em>how fast teams act</em> once they patch, whereas the survival curve shows <em>how long vulnerabilities truly last</em>.</p>
<h2 id="differencebetweensurvivalanalysisandmeanmedian">Difference Between Survival Analysis and Mean/Median</h2>
<p>Wait, we said that survival analysis is better to analyze time to patch to avoid the impact of outliers. But in this example, mean/median and survival analysis provide similar results. What is the added value? The reason is simple: we don’t have outliers in our production environments since our patching process is fully automated and effective.</p>
<p>To demonstrate the impact on heterogeneous data, we’ll use an outdated example from a non-production environment that lacks automated patching. </p>
<p>ESQL query:</p>
<pre><code>FROM qualys_vmdr.vulnerability_6months
  | WHERE elastic.environment == "my-outdated-non-production-environment"
  | WHERE qualys_vmdr.asset_host_detection.vulnerability.is_ignored == FALSE
  | EVAL vulnerability_age = DATE_DIFF("day", qualys_vmdr.asset_host_detection.vulnerability.first_found_datetime, qualys_vmdr.asset_host_detection.vulnerability.last_found_datetime)
  | STATS
      count=COUNT(*),
      count_closed_only=COUNT(*) WHERE qualys_vmdr.asset_host_detection.vulnerability.status == "Fixed",
      mean_observed_so_far=MEDIAN(vulnerability_age),
      mean_closed_only=MEDIAN(vulnerability_age) WHERE qualys_vmdr.asset_host_detection.vulnerability.status == "Fixed",
      median_observed_so_far=MEDIAN(vulnerability_age),
      median_closed_only=MEDIAN(vulnerability_age) WHERE qualys_vmdr.asset_host_detection.vulnerability.status == "Fixed"
</code></pre>
<p>|  | Observed so far | Closed only |
| :---- | :---- | :---- |
| Count | 833 | 322 |
| Mean | 178.7 (days) | 163.8 (days) |
| Median | 61 (days) | 5 (days) |
| Median survival | 527 (days) | N/A |</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltffeaa1a1fe66c960/6a7d860db437700dbb4d4014/image1.png" alt="" /></p>
<p>In this example, using mean and median yield very different results. Choosing a single representative metric can be challenging and potentially misleading. The survival analysis graph accurately represents our effectiveness in addressing vulnerabilities within this environment. </p>
<h2 id="finalthoughts">Final Thoughts</h2>
<p>The benefits of using survival analysis come not only from more accurate measurement but also from the insights into the dynamics of patching behaviour, showing where bottlenecks occur, factors that affect patching velocity and whether it aligns with our SLO. From a technical integration perspective, the use of survival analysis as part of our operational workflows and reporting can be achieved with minimal additional changes to our current Elastic Stack setup: survival analysis can run on the same cadence as our patching cycle with the results being pushed back into Kibana for visualization. The definitive advantage is to pair our existing operational metrics with survival analysis for both long-term trends and short-term performance tracking.  </p>
<p>Looking forward, we’re experimenting with additional new metrics like <strong>Arrival Rate</strong>, <strong>Burndown Rate</strong>, and <strong>Escape Rate</strong> that give us a way to move toward a more dynamic understanding of how vulnerabilities are really handled. </p>
<p><strong>Arrival Rate</strong> is the measure of how quickly new vulnerabilities are entering the environment. Knowing that fifty new CVEs show up each month, for example, tells us what to expect in the workload before we even start measuring patches. So the arrival rate is a metric that does not necessarily inform about the backlog, but more about the pressure applied to the system.</p>
<p><strong>Burndown Rate</strong> (trend) shows the other half of the equation: how quickly vulnerabilities are being remediated relative to how fast they arrive. </p>
<p><strong>Escape Rate</strong> adds yet another dimension by focusing on vulnerabilities that slip past the points where they should have been contained. In our context, an escape is about CVEs that miss patching windows or exceed SLO thresholds. An elevated escape rate doesn’t just show that vulnerabilities exist but it also shows that the process designed to control them is failing, whether because patching cycles are too slow, automation processes are lacking, or compensating controls are not working as intended.</p>
<p>Together, the metrics create a better picture: arrival rate tells us how much new risk is being introduced; burndown trends show whether we are keeping pace with that pressure or being overwhelmed by it; escape rates expose where vulnerabilities persist despite planned controls.</p>
<p>[1]:An outlier in statistics is a data point that is very far from the central tendency (or far from the rest of the values in a dataset). For example, if most vulnerabilities are patched within 30 days, but one takes 600 days, that 600-day case is an outlier. Outliers can pull averages upward or downward in ways that don’t reflect the “typical” experience. In the patching context, these are the especially slow-to-patch vulnerabilities that sit open far longer than the norm. They may represent rare but important situations, like systems that can’t be easily updated, or patches that require extensive testing. </p>
<p>[2]: Note: The current 6-month dataset includes both all vulnerabilities that remain open at the end of the observation period (independent of how long ago they have been open /first seen) and all vulnerabilities that were closed during the 6-month window. Despite this mixed cohort approach, survival curves from prior observation windows show consistent trends, particularly in the early part of the curve. The shape and slope over the first 30–60 days have proven remarkably stable across snapshots, suggesting that metrics like median time-to-patch and early-stage remediation behavior are not artifacts of the short observation window. While long-term estimates (e.g. 90th percentile) remain incomplete in shorter snapshots, the conclusions drawn from these cohorts still reflect persistent and reliable patching dynamics.</p>
<p>[3]:We kept the CDF on a monthly cadence for operational reporting (throughput and SLO adherence for work completed during the current month), while the Kaplan-Meier uses a 6-month window to properly handle censoring and expose tail risk across the broader cohort.</p>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/time-to-patch-metrics</link>
    <guid isPermaLink="false">time-to-patch-metrics</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Laura Voicu,Clement Fouque]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt31761e7b9110cf67/6a7d86106693f8a97d66111a/Security_Labs_Images_7.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 22 Oct 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Now in beta: New Detection as Code capabilities]]></title>
    <description><![CDATA[]]></description>
    <content:encoded><![CDATA[<p>Exciting news! Our Detections as Code (DaC) improvements to the <a href="https://github.com/elastic/detection-rules">detection-rules</a> repo are now in beta. In May this year, we shared the Alpha stages of our research into <a href="https://www.elastic.co/blog/detections-as-code-elastic-security">Rolling your own Detections as Code with Elastic Security</a>. Elastic is working on supporting DaC in Elastic Security. While in the future DaC will be integrated within the UI, the current updates are focused on the detection rules repo on main to allow users to set up DaC quickly and get immediate value with available tests and commands integration with Elastic Security. We have a considerable amount of <a href="https://dac-reference.readthedocs.io/en/latest/index.html">documentation</a> and <a href="https://dac-reference.readthedocs.io/en/latest/etoe_reference_example.html">examples</a>, but let’s take a quick look at what this means for our users.  </p>
<h2 id="whydac">Why DaC?</h2>
<p>From validation and automation to enhancing cross-vendor content, there are several reasons <a href="https://www.elastic.co/blog/detections-as-code-elastic-security#why-detections-as-code">previously discussed</a> to use a DaC approach for rule management. Our team of detection engineers have been using the detection rules repo for testing and validation of our rules for some time. We now can provide the same testing and validation that we perform in a more accessible way. We aim to empower our users by adding straightforward CLI commands within our detection-rules repo, to help manage rules across the full rule lifecycle between version control systems (VCS) and Kibana. This allows users to move, unit test, and validate their rules in a single command easily using CI/CD pipelines.</p>
<h2 id="improvingprocessmaturity">Improving Process Maturity</h2>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta126376ee8d18fc0/6a7d7e7a63e9594d0273ae16/image10.png" alt="" /></p>
<p>Security organizations are facing the same bottomline, which is that we can’t rely on static out-of-the-box signatures. At its core, DaC is a methodology that applies software development practices to the creation and management of security detection rules, enabling automation, version control, testing, and collaboration in the development &amp; deployment of security detections. Unit testing, peer review, and CI/CD enable software developers to be confident in their processes. These help catch errors and inefficiencies before they impact their customers. The same should be true in detection engineering. Fitting with this declaration here are some examples of some of the new features we are supporting. See our <a href="https://dac-reference.readthedocs.io/en/latest/">DaC Reference Guide</a> for complete documentation.</p>
<h3 id="bulkimportandexportofcustomrules">Bulk Import and Export of Custom Rules</h3>
<p>Custom rules can now be moved in bulk to and from Kibana using the <code>kibana import-rules</code> and <code>kibana export-rules</code> commands. Additionally, one can move them in bulk to and from TOML format to ndjson using the <code>import-rules-to-repo</code> and <code>export-rules-from-repo</code> commands. In addition to rules, these commands support moving exceptions and exception lists using the appropriate flag. The ndjson approach's benefit is that it allows engineers to manage and share a collection of rules in a single file (exported by the CLI or from Kibana), which is helpful when access is not permitted to the other Elastic environment. When moving rules using either of these methods, the rules pass through schema validation unless otherwise specified to ensure that the rules contain the appropriate data fields. For more information on these commands, please see the <a href="https://github.com/elastic/detection-rules/blob/DAC-feature/CLI.md"><code>CLI.md</code></a> file in detection rules. </p>
<h3 id="configurableunittestsvalidationandschemas">Configurable Unit Tests, Validation, and Schemas</h3>
<p>With this new feature, we've now included the ability to configure the behavior of unit tests and schema validation using configuration files. In these files, you can now set specific tests to be bypassed, specify only specific tests to run, and likewise with schema validation against specific rules. You can run this validation and unit tests at any time by running <code>make test</code>. Furthermore, you can now bring your schema (JSON file) to our validation process. You can also specify which schemas to use against which target versions of your Stack. For example, if you have custom schemas that only apply to rules in 8.14 while you have a different schema that should be used for 8.10, this can now be managed via a configuration file. For more information, please see our <a href="https://github.com/elastic/detection-rules/blob/DAC-feature/detection_rules/etc/_config.yaml">example configuration file</a> or use our <code>custom-rules setup-config</code> command from the detection rules repo to generate an example for you.</p>
<h3 id="customversioncontrol">Custom Version Control</h3>
<p>We now are providing the ability to manage custom rules using the same version lock logic that Elastic’s internal team uses to manage our rules for release. This is done through a version lock file that checks the hash of the rule contents and determines whether or not they have changed. Additionally, we are providing a configuration option to disable this version lock file to allow users to use an alternative means of version control such as using a git repo directly. For more information please see the <a href="https://dac-reference.readthedocs.io/en/latest/internals_of_the_detection_rules_repo.html#rule-versioning">version control section</a> of our documentation. Note that you can still rely on Kibana’s versioning fields.</p>
<p>Having these systems in place provides auditable evidence for maintaining security rules. Adopting some or all of these best practices can dramatically improve quality in maintaining and developing security rules.</p>
<h3 id="broaderadoptionofautomation">Broader Adoption of Automation</h3>
<p>While quality is critical, security teams and organizations face  growing rule sets to respond to an ever-expanding threat landscape. As such, it is just as crucial to reduce the strain on security analysts by providing rapid deployment and execution. For our repo, we have a single-stop shop where you can set your configuration, focus on rule development, and let the automation handle the rest.  </p>
<h4 id="loweringthebarriertoentry">Lowering the Barrier to Entry</h4>
<p>To start, simply clone or fork our detection rules repo, run <code>custom-rules setup-config</code> to generate an initial config, and import your rules. From here, you now have unit tests and validation ready for use. If you are using GitLab, you can quickly create CI/CD to push the latest rules to Kibana and run these tests. Here is an <a href="https://dac-reference.readthedocs.io/en/latest/core_component_syncing_rules_and_data_from_vcs_to_elastic_security.html#option-1-push-on-merge">example</a> of what that could look like:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9f4d0831347a9525/6a7d7e7d227b1c2c8859578e/image2.png" alt="Example CI/CD Workflow" title="Example CI/CD Workflow" /></p>
<h3 id="highflexibility">High Flexibility</h3>
<p>While we use GitHub CI/CD for managing our release actions, by no means are we prescribing that this is the only way to manage detection rules. Our CLI commands have no dependencies outside of their python requirements. Perhaps you have already started implementing some DaC practices, and you may be looking to take advantage of the Python libraries we provide. Whatever the case may be, we want to encourage you to try adopting DaC principles in your workflows and we would like to provide flexible tooling to accomplish these goals. </p>
<p>To illustrate an example, let’s say we have an organization that is already managing their own rules with a VCS and has built automation to move rules back and forth from deployment environments. However, they would like to augment these movements with testing based on telemetry which they are collecting and storing in a database. Our DaC features already provide custom unit testing classes that can run per rule. Realizing this goal may be as simple as forking the detection rules repo and writing a single unit test. The figure below shows an example of what this could look like.  </p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt98b581011c0ce777/6a7d7e8063e959670d73ae1a/image3.png" alt="Testing and Tuning via Data Source Input Workflow" title="Testing and Tuning via Data Source Input Workflow" /></p>
<p>This new unit test could utilize our unit test classes and rule loading to provide scaffolding to load rules from a file or Kibana instance. Next, one could create different integration tests against each rule ID to see if they pass the organization's desired results (e.g. does the rule identify the correct behaviors). If they do, the CI/CD tooling can proceed as originally planned. If they fail, one can use DaC tooling to move those rules to a “needs tuning” folder and/or upload those rules to a “Tuning” Kibana space. In this way, one could use a hybrid of our tooling and one's own tooling to keep an up to date Kibana space (or VCS controlled folder) of what rules require updates. As updates are made and issues addressed, they could also be continually synchronized across spaces, leading to a more cohesive environment.</p>
<p>This is just one idea of how one can take advantage of our new DaC features in your environment. In practice, there are a vast number of different ways they can be utilized.</p>
<h2 id="inpractice">In Practice</h2>
<p>Now, let’s take a look at how we can tie these new features together into a cohesive DaC strategy. As a reminder, this is not prescriptive. Rather, this should be thought of as an optional, introductory strategy that can be built on to achieve your DaC goals.</p>
<h3 id="establishingadacbaseline">Establishing a DaC Baseline</h3>
<p>In detection engineering, we would like collaboration to be a default rather than an exception. Detection Rules is a public repo precisely with this precept in mind. Now, it can become a basis for the community and teammates to not only collaborate with us, but also with each other. Let’s use the chart below as an example for what this could look like. </p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc51602caae5acb1a/6a7d7e82dd26d279fd2a7219/image1.png" alt="DaC Baseline Workflow" title="DaC Baseline Workflow" /></p>
<p>Reading from left to right, we have initial planning and prioritization and the subsequent threat research that drives the detection engineering. This process will look quite different for each user so we are not going to spend much time describing it here. However, the outcome will largely be similar, the creation of new detection rules. These could be in various forms like Sigma rules (more in a later blog), Elastic TOML rule files, or creating the rules directly in Kibana. Regardless of format, once created these rules need to be staged. This would either occur in Kibana, your VCS, or both. From a DaC perspective, the goal is to sync the rules such that the process/automation are aware of these new additions. Furthermore, this provides the opportunity for peer review of these additions — the first stage of collaboration. </p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0bab5165d38f1a3b/6a7d7e8563e9596da273ae1e/image8.png" alt="Peer Review Workflow" title="Peer Review Workflow" /></p>
<p>This will likely happen in your version control system; for instance, in GitHub one could use a PR with required approvals before merging back into a main branch that acts as the authoritative source of reviewed rules. The next step is for testing and validation, this step could additionally occur before peer review and this is largely up to the desired implementation. </p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd9538c64ce01f023/6a7d7e8896b5a6cb1d8785af/image11.png" alt="Validation to Production Workflow" title="Validation to Production Workflow" /></p>
<p>In addition to any other internal release processes, by adhering to this workflow, we can reduce the risk of malformed rules and errant mistakes from reaching both our customers and the community. Additionally, having the evidence artifacts, passing unit tests, schema validation, etc., inspires confidence and provides control for each user to choose what risks they are willing to accept. </p>
<p>Once deployed and distributed, rule performance can be monitored from Kibana. Updates to these rules can be made either directly from Kibana or through the VCS. This will largely be dependent on the implementation specifics, but in either case, these can be treated very similarly to new rules and pass through the same peer review, testing, and validation processes.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt513a1ecdf3ad304d/6a7d7e8b77b0347a2c3fc5ce/image14.png" alt="Tuning Production Deployment Workflow" title="Tuning Production Deployment Workflow" /></p>
<p>As shown in the figure above, this can provide a unified method for handling rule updates whether from the community, customers, or from internal feedback. Since the rules ultimately exist as version-controlled files, there is a dedicated format source of truth to merge and test against. </p>
<p>In addition to the process quality improvements, having authoritative known states can empower additional automation. As an example, different customers may require different testing or perhaps different data sources. Instead of having to parse the rules manually, we provide a unified configuration experience where users can simply bring their own config and schemas and be confident that their specific requirements are met. All of this can be managed automatically via CI/CD. With a fully automated DaC setup, one can take advantage of this system entirely from VCS and Kibana without needing to write additional code. Let’s take a look at an example of what this could look like. </p>
<h3 id="example">Example</h3>
<p>For this example, we are going to be acting as an organization that has 2 Kibana spaces they want to manage via DaC. The first is a development space that rule authors will be using to write detection rules (so let’s assume there are some preexisting rules already available). There will also be some developers that are writing detection rules directly in TOML file formats and adding them to our VCS, so we will need to manage synchronization of these. Additionally, this organization wants to enforce unit testing and schema validation with the option for peer review on rules that will be deployed to a production space in the same Kibana instance. Finally, the organization wants all of this to occur in an automated manner with no requirement to either clone detection rules locally or write rules outside of a GUI. </p>
<p>In order to accomplish this we will need to make use of a few of the new DaC features in detection rules and write some simple CI/CD workflows. In this example we are going to be using GitHub. Additionally, you can find a video walkthrough of this example <a href="https://dac-reference.readthedocs.io/en/latest/etoe_reference_example.html#demo-video">here</a>. As a note, if you wish to follow along you will need to fork the detection rules repo and create an initial configuration using our <code>custom-rules setup-config</code> command. Also for general step by step instructions on how to use the DAC features, see this <a href="https://dac-reference.readthedocs.io/en/latest/etoe_reference_example.html#quick-start-example-detection-rules-cli-commands">quickstart guide</a>, which has several example commands.</p>
<h4 id="developmentspacerulesynchronization">Development Space Rule Synchronization</h4>
<p>First we are going to synchronize from Kibana -&gt; GitHub (VCS). To do this we will be using the <code>kibana import-rules</code> and <code>kibana export-rules</code> detection rules commands. Additionally, in order to keep the rule versions synchronized we will be using the locked versions file as we are wanting both our VCS and Kibana to be able to overwrite each other with the latest versions. This is not required for this setup, either Kibana or GitHub (VCS) could be used authoritatively instead of the locked versions file. But we will be using it for convenience. </p>
<p>The first step is for us to make a manual dispatch trigger that will pull the latest rules from Kibana upon request. In our setup this could be done automatically; however, we want to give rule authors control for when they want to move their rules to the VCS as the development space in Kibana is actively used for development and the presence of a new rule does not necessarily mean the rule is ready for VCS. The manual dispatch section could look like the following <a href="https://dac-reference.readthedocs.io/en/latest/core_component_syncing_rules_and_data_from_elastic_security_to_vcs.html#option-1-manual-dispatch-pull">example</a>:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc93fb309b116d124/6a7d7e8e42a117b06a959085/image15.png" alt="" /></p>
<p>With this trigger in place, we now can write 4 additional jobs that will trigger on this workflow dispatch. </p>
<ol>
<li>Pull the rules from the desired Kibana space. </li>
<li>Update the version lock file. </li>
<li>Create a PR request for review to merge into the main branch in GitHub. </li>
<li>Set the correct target for the PR.</li>
</ol>
<p>These jobs could look like this also from the same <a href="https://dac-reference.readthedocs.io/en/latest/core_component_syncing_rules_and_data_from_elastic_security_to_vcs.html#option-1-manual-dispatch-pull">example</a>:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt78dfc1ef5950d471/6a7d7e911967eaf9fc32d833/image12.png" alt="" /></p>
<p>Now, once we run this workflow we should expect to see a PR open with the new rules from the Kibana Dev space. We also need to synchronize rules from GitHub (VCS) to Kibana. For this we will need to create a triggers on pull request:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9756450ee83d43cc/6a7d7e9451156a8a672bf7f5/image4.png" alt="" /></p>
<p>Next, we just need to create a job that uses the <code>kibana import-rules</code> command to push the rule files from the given PR to Kibana. See the second <a href="https://dac-reference.readthedocs.io/en/latest/core_component_syncing_rules_and_data_from_vcs_to_elastic_security.html#option-1-push-on-merge">example</a> for the complete workflow file.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt586a6a4f65b99f24/6a7d7e97fc63ab0eac649f29/image5.png" alt="" /></p>
<p>With these two workflows complete we now have synchronization of rules between GitHub and the Kibana Dev space. </p>
<h3 id="productionspacedeployment">Production Space Deployment</h3>
<p>With the Dev space synchronized, now we need to handle the prod space. As a reminder, for this we need to enforce unit testing, schema validation, available peer review for PRs to main, and on merge to main auto push to the prod space. To accomplish this we will need two workflow files. The first will run unit tests on all pull requests and pushes to versioned branches. The second will push the latest rules merged to main to the prod space in Kibana. </p>
<p>The first workflow file is very simple. It has an on push and pull_request trigger and has the core job of running the <code>test</code> command shown below. See this <a href="https://dac-reference.readthedocs.io/en/latest/core_component_syncing_rules_and_data_from_elastic_security_to_vcs.html#sub-component-3-optional-unit-testing-rules-via-ci-cd">example</a> for the full workflow.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt586a6a4f65b99f24/6a7d7e97fc63ab0eac649f29/image5.png" alt="" /></p>
<p>With this <code>test</code> command we are performing unit tests and schema validation with the parameters specified in our config files on all of our custom rules. Now we just need the workflow to push the latest rules to the prod space. The core of this workflow is the <code>kibana import-rules</code>command again just using the prod space as the destination. However, there are a number of additional options provided to this workflow that are not necessary but nice to have in this example, such as options to overwrite and update exceptions/exception lists as well as rules. The core job is shown below. Please see <a href="https://dac-reference.readthedocs.io/en/latest/core_component_syncing_rules_and_data_from_vcs_to_elastic_security.html#option-1-push-on-merge">this example</a> for the full workflow file.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte0597b33f3d56ec5/6a7d7e9ac2cc0950852465f3/image7.png" alt="" /></p>
<p>And there we have it, with those 4 workflow files we have a synchronized development space with rules passing through unit testing and schema validation. We have the option for peer review through the use of pull requests, which can be made as requirements in GitHub before allowing for merges to main. On merge to main in GitHub we also have an automated push to the Kibana prod space, establishing our baseline of rules that have passed our organizations requirements and are ready for use. All of this was accomplished without writing additional Python code, just by using our new DaC features in GitHub workflows.</p>
<h2 id="conclusion">Conclusion</h2>
<p>Now that we’ve reached this milestone, you may be wondering what’s next? We’re planning to spend the next few cycles continuing to test edge cases and incorporating feedback from the community as part of our business-as-usual sprints. We also have a backlog of features request considerations so if you want to voice your opinion, checkout the issues titled <code>[FR][DAC] Consideration:</code> or open a similar new issue if it’s not already recorded. This will help us to prioritize the most important features for the community.</p>
<p>We’re always interested in hearing use cases and workflows like these, so as always, reach out to us via <a href="https://github.com/elastic/detection-rules/issues">GitHub issues</a>, chat with us in our <a href="https://elasticstack.slack.com/archives/C06TE19EP09">security-rules-dac</a> slack channel, and ask questions in our <a href="https://discuss.elastic.co/c/security/endpoint-security/80">Discuss forums</a>!</p>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/dac-beta-release</link>
    <guid isPermaLink="false">dac-beta-release</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Mika Ayenson,Eric Forte]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0ad39f15bb51a731/6a7d7e9cead8ec7af3ba7ac1/Security_Labs_Images_18.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 08 Aug 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Unveiling malware behavior trends]]></title>
    <description><![CDATA[An analysis of a diverse dataset of Windows malware extracted from more than 100,000 samples revealing insights into the most prevalent tactics, techniques, and procedures.]]></description>
    <content:encoded><![CDATA[<h2 id="preamble">Preamble</h2>
<p>When prioritizing detection engineering efforts, it's essential to understand the most prevalent tactics, techniques, and procedures (TTPs) observed in the wild. This knowledge helps defenders make informed decisions about the most effective strategies to implement - especially where to focus engineering efforts and finite resources.</p>
<p>To highlight these prevalent TTPs, we analyzed over <a href="https://gist.github.com/Samirbous/eebeb8f776f7ab2d51cdd2ac05669dcf">100,000 Windows malware samples</a> extracted over several months from one of our dynamic malware analysis tools, <a href="https://www.elastic.co/security-labs/click-click-boom-automating-protections-testing-with-detonate">Detonate</a>. To generate this data and alerts, we leveraged Elastic Defend behavior (mapped to MITRE ATT&amp;CK) and <a href="https://www.elastic.co/guide/en/security/current/configure-endpoint-integration-policy.html#memory-protection">memory threat detection</a> rules. It should be noted that this dataset is not exhaustive, it may not represent the entire spectrum of malware behavior, and specifically does not include long-term or interactive activity.</p>
<p>Below an <a href="https://www.elastic.co/blog/esql-elasticsearch-piped-query-language">ES|QL</a> query to summarize our dataset by file type:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd33f907667e45ce4/6a7d868d1967ea043b32d938/image12.png" alt="Dataset by extension - 20 unique file types" title="Dataset by extension - 20 unique file types" /></p>
<h2 id="tactics">Tactics</h2>
<p>Beginning with tactics, we aggregated the alerts generated by this corpus of malware samples and organized them according to the counts of <a href="https://www.elastic.co/guide/en/ecs/current/ecs-process.html#field-process-entity-id"><code>process.entity_id</code></a> and alerts. As depicted in the image below, the most frequent tactics included defense evasion, privilege escalation, execution, and persistence. Certain tactics commonly linked with post-exploitation activities, such as lateral movement, provided an anticipated lower prevalence because these actions are commonly manually driven by the threat actor after the initial implant is established vs. being automated by the malware in our dataset.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5ca41e66a1715602/6a7d8690ead8ec1f8bba7bde/image9.png" alt="Tactics by volume" title="Tactics by volume" /></p>
<p>In the following sections, we will delve into each tactic and the techniques and sub-techniques of each that exerted the most influence.</p>
<h3 id="defenseevasion">Defense Evasion</h3>
<p>Defense Evasion involves methods employed by adversaries to avoid detection by security teams or capabilities. The foremost tactic detected was defense evasion, triggering 189 distinct detection rules (nearly 40% of our current Windows rules). The primary techniques noted are associated with <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1055%22&amp;type=code">code injection</a>, <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22Impair+Defenses%22&amp;type=code">defense tampering</a>, <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22Masquerade+Task+or+Service%22&amp;type=code">masquerading</a>, and <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1218%22&amp;type=code">system binary proxy execution</a>.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf317b287f82ecbb9/6a7d8693de2315ac76fd4ee3/image15.png" alt="Top observed defense evasion techniques" title="Top observed defense evasion techniques" /></p>
<p>When we pivot by sub-techniques, it becomes evident that certain advanced techniques such as <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22DLL+Side-Loading%22&amp;type=code&amp;p=1">DLL side-loading</a> and <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts%20%22Parent%20PID%20Spoofing%22&amp;type=code">Parent PID Spoofing</a> have become increasingly popular, even among non-targeted malwares. Both are frequently linked with code injection and masquerading.</p>
<p>Furthermore, system binary proxies <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1218.011%22&amp;type=code"><code>Rundll32</code></a> and <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1218.010%22&amp;type=code"><code>Regsvr32</code></a> remain highly abused, with a notable rise in the utilization of malicious <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1218.007%22&amp;type=code">MSI installers</a> for malware delivery. The practice of <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22Masquerade+Task+or+Service%22&amp;type=code">masquerading</a> as legitimate system binaries, whether through renaming or process hollowing, remains prevalent as well, serving as a means to evade user suspicion.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5fbcda6cfdd805db/6a7d86965967e5e0405da5db/image6.png" alt="Top observed defense evasion sub-techniques" title="Top observed defense evasion sub-techniques" /></p>
<p>Tampering with Windows Defender stands out as the most frequently observed defense evasion tactic, emphasizing the importance for defenders to acknowledge that adversaries will attempt to obscure their activities. </p>
<p>Process Injection is prevalent across various malware families, whether they target legitimate system binaries remotely to blend in or employ self-injection (sometimes paired with DLL side-loading through a trusted binary). Furthermore, there is a noticeable uptick in the use of NTDLL unhooking to bypass security solutions reliant on user-mode APIs monitoring (Elastic Defend is not impacted).</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5698a98a0cf35f91/6a7d8699e02fac67ae5d359e/image16.png" alt="The most effective endpoint behavior rules for defense evasion" title="The most effective endpoint behavior rules for defense evasion" /></p>
<p>From our shellcode alerts we can clearly see that self-injection is more prevalent than remote: </p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdf4e4753659f9b6d/6a7d869c448e4e58d45bdbd6/image7.png" alt="Shellcode alerts volume by infection target type (local vs remote)" title="Shellcode alerts volume by infection target type (local vs remote)" /></p>
<p>Almost 50 unique vendors’ binaries abused for DLL side-loading, of which Microsoft is the top choice: </p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9bcf27fbfcb0e930/6a7d869f05b7b5d394188bfb/image19.png" alt="DLL side-load by host process code signature subject name" title="DLL side-load by host process code signature subject name" /></p>
<p>Defense evasion comprises various techniques and sub-techniques necessitating comprehensive coverage due to their frequent occurrence. For instance, apart from <a href="https://www.elastic.co/guide/en/security/current/configure-endpoint-integration-policy.html#memory-protection">memory threat protection</a>, <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts++name+%3D+%22Defense+Evasion%22&amp;type=code">half</a> of our rules are specifically tailored to address this tactic.</p>
<h3 id="privilegeescalation">Privilege Escalation</h3>
<p>This tactic consists of techniques that adversaries use to gain greater permissions on a system or network. The most commonly used techniques relate to <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1134%22&amp;type=code">access token manipulation</a>, execution through privileged <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1543.003%22&amp;type=code">system services</a>, and bypassing <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1548.002%22&amp;type=code">User Account Control</a>.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb1e54ec1c73d253e/6a7d86a11967ead38432d93c/image10.png" alt="Privilege escalation techniques observed in the dataset" title="Privilege escalation techniques observed in the dataset" /></p>
<p>The most frequently observed sub-technique involved impersonation as the Trusted Installer service, which aligns closely with defense evasion and often precedes attempts to manipulate system-protected resources. </p>
<p>Concerning User Account Control bypass, the primary method we observed was elevation by <a href="https://medium.com/tenable-techblog/uac-bypass-by-mocking-trusted-directories-24a96675f6e">mimicking trusted directories</a>, which is also related to DLL side-loading. Additionally, other methods like elevation via <a href="https://github.com/decoder-it/psgetsystem">extended startupinfo</a> (elevated parent PID spoofing) are increasingly prevalent among commodity malware.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbc66ee7f379161c5/6a7d86a4e3a219368e99c7b4/image1.png" alt="Privilege escalation top observed sub-techniques" title="Privilege escalation top observed sub-techniques" /></p>
<p>As evident from the list below, there's a notable rise in the use of <a href="https://www.elastic.co/security-labs/stopping-vulnerable-driver-attacks">vulnerable drivers</a> (BYOVD) to manipulate protected objects and acquire kernel mode execution privileges. </p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5c1a44230ace6c8a/6a7d86a7dd26d25eee2a72e9/image2.png" alt="The most effective endpoint behavior rules for privilege escalation" title="The most effective endpoint behavior rules for privilege escalation" /></p>
<p>Below, you'll find a list of the most commonly exploited drivers triggered by our <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts%20vulndriver&amp;type=code">YARA rules</a>:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4ff6465fd24d840e/6a7d86aa5967e523f55da5df/image18.png" alt="Top triggered yara rules for vulnerable drivers detection" title="Top triggered yara rules for vulnerable drivers detection" /></p>
<h3 id="execution">Execution</h3>
<p>Execution encompasses methods that lead to running adversary-controlled code on a local or remote system. These techniques are frequently combined with methods from other tactics to accomplish broader objectives, such as network reconnaissance or data theft. </p>
<p>The most common techniques observed here involved <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22Command+and+Scripting+Interpreter%22+%5B%22windows%22%5D&amp;type=code">Windows command and scripting languages</a>, with the proxying of execution via the <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1047%22&amp;type=code">Windows Management Instrumentation</a> (WMI) interface closely trailing behind.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt04c29bb44eab1d22/6a7d86ad227b1c70715958b8/image21.png" alt="Execution techniques observed in our dataset" title="Execution techniques observed in our dataset" /></p>
<p><a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1059.001%22&amp;type=code">Powershell</a> remains a preferred scripting language for malware execution chains, followed by <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1059.007%22&amp;type=code">Javascript</a> and <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1059.005%22&amp;type=code">VBscript</a>. Multi-stage malware delivery routinely involves a combination of two or more scripting languages.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt72bf1af35e8ae57c/6a7d86b0227b1c55eb5958bc/image5.png" alt="Execution top observed sub-techniques" title="Execution top observed sub-techniques" /></p>
<p>Here is a list of the most frequently triggered endpoint behavior detections for this tactic:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbc15902eaaaa4b70/6a7d86b36693f812f866112e/image13.png" alt="Frequently triggered execution detections" title="Frequently triggered execution detections" /></p>
<p>Windows' default scripting languages remain the top preference for malware execution. However, there has been a slight uptick in the shift towards using other third-party scripting interpreters like Python, AutoIt, Java and Lua.</p>
<h3 id="persistence">Persistence</h3>
<p>It's common for malware to install itself on an infected host. No surprises here: the most frequently observed persistence methods include <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1053.005%22&amp;type=code">scheduled tasks</a>, the <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1547.001%22&amp;type=code">run key and startup folder</a>, and <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1543.003%22&amp;type=code">Windows services</a> (which typically require administrator privileges).</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4c7825d5fd471d1a/6a7d86b677b034a3083fc6fd/image14.png" alt="Top observed sub-techniques for persistence" title="Top observed sub-techniques for persistence" /></p>
<p>The top three persistence sub-techniques depicted in the list below are also commonly encountered in regular software installations. Therefore, it's necessary to dissect them into multiple detections with additional suspicious signals to reduce false positives and enhance precision.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5ca419988ce4b04b/6a7d86b8e02fac48705d35a6/image8.png" alt="Top triggered alerts for persistence" title="Top triggered alerts for persistence" /></p>
<h3 id="initialaccess">Initial Access</h3>
<p>Considering the dataset's composition, initial access was associated with primarily macro-enabled documents and Windows shortcut objects. Although a significant portion of the detonated samples also involved other formats, such as ISO/VHD containers with MSI installers extensively utilized for delivery, their genuine malicious behavior typically manifests in areas such as defense evasion and persistence.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt00fcd9649a225202/6a7d86bbb43770659d4d4024/image17.png" alt="Top sub-techniques for initial access" title="Top sub-techniques for initial access" /></p>
<p>The most frequently abused Microsoft-signed binaries originating from malicious Microsoft Office documents align closely with execution and defense evasion tactics, command and scripting interpreters, and system binary proxy execution.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte28a834992f439f2/6a7d86bfe02facdb5e5d35aa/image11.png" alt="Top spawned child processes from malicious office documents" title="Top spawned child processes from malicious office documents" /></p>
<p>Here is a list of the most frequently triggered detections for initial access, regarding <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1566.001%22&amp;type=code">phishing attachments</a>:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc83fb53771308bae/6a7d86c2fc63ab927864a052/image4.png" alt="Top triggered rules for initial access via malicious attachments" title="Top triggered rules for initial access via malicious attachments" /></p>
<h3 id="credentialaccess">Credential Access</h3>
<p>Credential access in malware is frequently linked to information stealers. The most targeted credentials are typically associated with <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1555.004%22&amp;type=code">Windows Credential Manager</a> and <a href="https://github.com/search?q=repo%3Aelastic%2Fprotections-artifacts+%22T1555.003%22&amp;type=code">browser password</a> stores. Domain and system-protected credentials require elevated privileges and are more likely a feature of a subsequent stage.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3bec30a102170c42/6a7d86c5448e4e7abc5bdbda/image20.png" alt="Top observed credential access sub-techniques" title="Top observed credential access sub-techniques" /></p>
<p>Below a breakdown of the endpoint behavior detections that triggered the most on credentials access: </p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt914e911c4b70b725/6a7d86c863e95953ef73af3f/image3.png" alt="Frequently triggered credential access-related detection rules " title="Frequently triggered credential access-related detection rules" /></p>
<p>The majority of credentials access behaviors resemble typical file access events. Therefore, it's essential to correlate and enrich them with additional signals to reduce false positives and enhance comprehension.</p>
<h2 id="conclusion">Conclusion</h2>
<p>Even though this small dataset of about <a href="https://gist.github.com/Samirbous/eebeb8f776f7ab2d51cdd2ac05669dcf">100,000 malware samples</a> represents only a fraction of the possible malware in the wild right now, we can still derive important insights from it about the most common TTPs using our behavioral detections. Those insights help us make decisions about detection engineering priorities, and defenders should make that part of their strategies.</p>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/unveiling-malware-behavior-trends</link>
    <guid isPermaLink="false">unveiling-malware-behavior-trends</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Samir Bousseaden]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc80e1b5c35bec0f5/6a7d86cbead8ec3636ba7bee/Security_Labs_Images_20.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 20 Mar 2024 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elastic Global Threat Report Multipart Series Overview]]></title>
    <description><![CDATA[Each month, the Elastic Security Labs team dissects a different trend or correlation from the Elastic Global Threat Report. This post provides an overview of those individual publications.]]></description>
    <content:encoded><![CDATA[<p>When we <a href="https://www.elastic.co/security-labs/2022-elastic-global-threat-report-announcement">announced</a> the inaugural Elastic Global Threat Report last year, the Elastic Security Labs team knew we wanted to follow it with a series that went a little deeper on several topics like trends and forecasting. Not only would this allow us to keep the report concise, but it would provide us with a way to be more transparent by diving deep.</p>
<p>This post will be updated with each new article in the series, published monthly:</p>
<ul>
<li><a href="https://www.elastic.co/blog/elastic-global-threat-report-breakdown-defense-evasion">Topic: Defense Evasion</a></li>
<li><a href="https://www.elastic.co/blog/elastic-global-threat-report-breakdown-credential-access">Topic: Credential Access</a></li>
</ul>
<p>In April, we published an <a href="https://ela.st/gtr">updated</a> version of the Global Threat Report Spring Edition which included new insights online and interactive.</p>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/gtr-multipart-series-overview</link>
    <guid isPermaLink="false">gtr-multipart-series-overview</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Devon Kerr]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb71b18ea41269efc/6a7d8107ea068d1620f07227/gtr-blog-image-720x420.png" length="0" type="image/png"/>
    <pubDate>Mon, 24 Apr 2023 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elastic’s 2022 Global Threat Report: A roadmap for navigating today’s growing threatscape]]></title>
    <description><![CDATA[Threat intelligence resources like the 2022 Elastic Global Threat Report are critical to helping teams evaluate their organizational visibility, capabilities, and expertise in identifying and preventing cybersecurity threats.]]></description>
    <content:encoded><![CDATA[<p>Staying up-to-date on the current state of security and understanding the implications of today’s growing threat landscape is critical to my role as CISO at Elastic. Part of this includes closely following the latest security threat reports, highlighting trends, and offering valuable insights into methods bad actors use to compromise environments.</p>
<p>Threat intelligence resources like the <a href="https://www.elastic.co/explore/security-without-limits/global-threat-report">2022 Elastic Global Threat Report</a> are critical to helping my team evaluate our organization’s visibility, capabilities, and expertise in identifying and preventing cybersecurity threats. It helps us answer questions such as:</p>
<ul>
<li>How is our environment impacted by the current and emerging threats identified in this report?</li>
<li>Does this new information change our risk profile and impact our risk analysis?</li>
<li>What adjustments do we need to make to our controls?</li>
<li>Are we lacking visibility in any areas?</li>
<li>Do we have the right detections in place?</li>
<li>How might these insights affect my team’s workflows?</li>
</ul>
<p>Elastic’s threat report provides a real-world roadmap to help my team make the connections necessary to strengthen our security posture. It influences our overall program roadmaps, helping us prioritize where we focus our resources, including adjusting our defenses, testing incident response plans, and identifying updates for our security operations center (SOC). And perhaps most importantly, the report underscores our belief that providing open, transparent, and accessible security for all organizations is key to defending ourselves against cybersecurity threats.</p>
<h2 id="checkyourcloudsecurityandthencheckitagain">Check your cloud security, and then check it again</h2>
<p>Threat reports often reinforce many of the existing trends and phenomena we see within security, but they can also reveal some unexpected insights. While the cloud enables organizations to operate faster and at scale, it also creates security gaps that leave room for potential attacks as threat actors continue shifting their focus to the cloud.</p>
<p>The Elastic Global Threat Report revealed that nearly 40% of all malware infections are on Linux endpoints, further emphasizing the need for better cloud security. With <a href="https://www.redhat.com/en/resources/state-of-linux-public-cloud-solutions-ebook">nine out of the top ten public clouds running on Linux</a>, this statistic is an important reminder to organizations not to rely solely on their cloud provider’s standard configurations for security.</p>
<p>The findings further revealed that approximately 57% of cloud security events were attributed to AWS, followed by 22% for Google Cloud and 21% for Azure, and that 1 out of every 3 (33%) cloud alerts was related to credential access across all cloud service providers.</p>
<p>While the data points to an increased need for organizations to properly secure their cloud environments, it also reinforces our belief that cloud security posture management (CSPM) needs to evolve similarly to endpoint security.</p>
<p>Initially, endpoint security relied on simple antivirus, which was only as good as its antivirus signatures. To avoid increasingly sophisticated malware and threats, endpoint security evolved by employing more advanced technologies like next-gen antivirus with machine learning and artificial intelligence. CSPM is currently facing a similar situation. Right now, we are closer to the bottom of the cloud security learning curve than the top, and our technologies and strategies must continue to evolve to manage new and emerging threats.</p>
<p>The Elastic Global Threat Report demonstrates that native tools and traditional security tactics are ineffective when implemented in cloud environments and offers recommendations for how organizations can adapt to the evolving threat landscape.</p>
<h2 id="getthebasicsrightfirst">Get the basics right first</h2>
<p>Security leaders and teams should leverage insights from this report to inform their priorities and adjust their workflows accordingly.</p>
<p>The findings clearly show why focusing on and improving basic security hygiene is so crucial to improved security outcomes. Too often, an organization’s environment is compromised by something as simple as a weak password or failure to update default configurations. Prioritizing security fundamentals — identity and access management, patching, threat modeling, password awareness, and multi-factor authentication — is a simple yet effective way for security teams to prevent and protect against potential threats.</p>
<h2 id="developsecurityintheopen">Develop security in the open</h2>
<p>Organizations should consider adopting an open approach to security. For example, Elastic’s threat report links to our recent publication of <a href="https://github.com/elastic/protections-artifacts">protection artifacts</a>, which transparently shares endpoint behavioral logic that we develop at Elastic to identify adversary tradecraft and make it freely available to our community.</p>
<p>The report also highlights how Elastic Security’s <a href="https://github.com/elastic/detection-rules/tree/main/rules">prebuilt detection rules</a> map to the MITRE ATT&amp;CK matrix for each cloud service provider. As an adopter of the MITRE framework since its inception, Elastic understands the importance of mapping detection rules to an industry standard. For my team, this helps us have deeper insights into the breadth and depth of our security posture.</p>
<p>Providing open detection rules, open artifacts, and open code enables organizations to focus on addressing gaps in their security technology stack and developing risk profiles for new and emerging threats. Without openness and transparency in security, organizations are putting themselves at greater risk of tomorrow’s cybersecurity threats.</p>
<p>Download the <a href="https://www.elastic.co/explore/security-without-limits/global-threat-report">2022 Elastic Global Threat Report</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/elastics-2022-global-threat-report-a-roadmap-for-navigating-todays-growing-threatscape</link>
    <guid isPermaLink="false">elastics-2022-global-threat-report-a-roadmap-for-navigating-todays-growing-threatscape</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Mandy Andress]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltf525dd6df010a3f7/6a7d7fbb1967ea3ccb32d873/gtr-blog-image-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 08 Dec 2022 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Forecast and Recommendations: 2022 Elastic Global Threat Report]]></title>
    <description><![CDATA[With the release of our first Global Threat Report at Elastic, customers, partners, and the security community at large are able to identify many of the focus areas our team has had over the past 12 months.]]></description>
    <content:encoded><![CDATA[<p>Today, we released our first-ever <a href="https://www.elastic.co/explore/security-without-limits/global-threat-report">Global Threat Report</a> at Elastic. Now, customers, partners, and the security community at large will be able to identify many of the focus areas our team has had over the past 12 months. In addition to a technical perspective, this report also brings along a series of strategic recommendations for executives and security leaders alike: a summarized, accurate perspective of where we can expect to see adversaries move over the coming months.</p>
<p>Our hope is that threat researchers and the security industry as a whole will use this report to prepare for the next set of threats and campaigns. At Elastic, we are ensuring that our customers using the Elastic Security solution are best protected from these types of threats, including endpoint and cloud capabilities for automated protection.</p>
<p>This year, our report included six key forecasts and recommendations for strategists and practitioners to stay better informed of potential directions that threat actors may focus on in 2023 and beyond. Below, we summarize the first three of our forecasts. Further details on these and our other recommendations are available in our full, <a href="https://www.elastic.co/explore/security-without-limits/global-threat-report">downloadable report</a> for 2022:</p>
<blockquote>
  <p>Adversaries will continue to abuse built-in binary proxies to evade security instrumentation. </p>
  <p>The use of proven adversarial tactics remains a key area of focus for observed threat groups, and this year remains no different. Hostile groups leverage legitimate, native system binaries to load malicious software — evading many detection strategies used by modern enterprises.</p>
  <p>With this continued focus, Elastic Security has enhanced our deep visibility and pre-built protections, including numerous rules and signatures, alongside ML models to detect these threats faster and more effectively.</p>
</blockquote>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt20ae933e37ebdeca/6a7d807305b7b5490e188b4a/blog-elastic-threat2022-1.jpg" alt="" /></p>
<blockquote>
  <p>LNK and ISO payloads will replace more conventional script and document payloads.</p>
  <p>Adversarial behavior focuses on finding easier, more efficient pathways for attack — and this year, it is no different. System defaults have forced threat groups to pivot their strategies to leverage LNK and ISO payloads over familiar scripts and documents we have observed in the past.</p>
  <p>LNK and ISO files are often used to smuggle malicious software into enterprises because most security technologies don't inspect them. Elastic Security has focused on building instrumentation into our products and platform, allowing us to determine the exact mechanisms used to better build a defense against these malicious acts.</p>
</blockquote>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd0d59d572b823326/6a7d8076de2315c202fd4e32/blog-elastic-threat2022-2.jpg" alt="" /></p>
<blockquote>
  <p>Valid IAM accounts will continue to be a target for adversaries.</p>
  <p>The early stages of many attacks focus on credential theft in all forms; however, IAM and administrative credentials often remain the area of focus for many adversarial groups looking to evade detection and avoid exploitation of services. </p>
  <p>Understanding standard account actions and user behaviors exhibited in environments is critical to defending them, and ensuring we have a comprehensive library of detections alongside integration capabilities within the stack has provided a strong foundation in detecting threats earlier.</p>
</blockquote>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb668bc2bad27ee58/6a7d8078437e0f8743dd85ba/blog-elastic-threat2022-3.jpg" alt="" /></p>
<p>This is just a small introduction to the findings found in the report. Far greater detail, recommendations, and source data are available in the <a href="https://www.elastic.co/explore/security-without-limits/global-threat-report">2022 Elastic Global Threat Report</a>.</p>
<p>Those looking to learn more about the threats we observed and the mechanisms adversarial groups leveraged over the last year can read far more detailed information in our full report — alongside many recommendations and findings we have leveraged to help shape the strategy used within the Elastic Security solution, and future feature roadmap.</p>
<p>Feel free to check out the full <a href="https://www.elastic.co/explore/security-without-limits/global-threat-report">2022 Elastic Global Threat Report</a> here.</p>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/forecast-and-recommendations-2022-elastic-global-threat-report</link>
    <guid isPermaLink="false">forecast-and-recommendations-2022-elastic-global-threat-report</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Santosh Krishnan]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte57e46ec17e7a8de/6a7d807b05b7b55d5d188b4e/gtr-blog-image-720x420.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 30 Nov 2022 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Getting the Most Out of Transformers in Elastic]]></title>
    <description><![CDATA[In this blog, we will briefly talk about how we fine-tuned a transformer model meant for a masked language modeling (MLM) task, to make it suitable for a classification task.]]></description>
    <content:encoded><![CDATA[<h2 id="preamble">Preamble</h2>
<p>In 8.3, our Elastic Stack Machine Learning team introduced a way to import <a href="https://www.elastic.co/guide/en/machine-learning/master/ml-nlp-model-ref.html">third party Natural Language Processing (NLP) models</a> into Elastic. As security researchers, we HAD to try it out on a security dataset. So we decided to build a model to identify malicious command lines by fine-tuning a pre-existing model available on the <a href="https://huggingface.co/models">Hugging Face model hub</a>.</p>
<p>Upon finding that the fine-tuned model was performing (surprisingly!) well, we wanted to see if it could replace or be combined with our previous <a href="https://www.elastic.co/blog/problemchild-detecting-living-off-the-land-attacks">tree-based model</a> for detecting Living off the Land (LotL) attacks. But first, we had to make sure that the throughput and latency of this new model were reasonable enough for real-time inference. This resulted in a series of experiments, the results of which we will detail in this blog.</p>
<p>In this blog, we will briefly talk about how we fine-tuned a transformer model meant for a masked language modeling (MLM) task, to make it suitable for a classification task. We will also look at how to import custom models into Elastic. Finally, we’ll dive into all the experiments we did around using the fine-tuned model for real-time inference.</p>
<h2 id="nlpforcommandlineclassification">NLP for command line classification</h2>
<p>Before you start building NLP models, it is important to understand whether an <a href="https://www.ibm.com/cloud/learn/natural-language-processing">NLP</a> model is even suitable for the task at hand. In our case, we wanted to classify command lines as being malicious or benign. Command lines are a set of commands provided by a user via the computer terminal. An example command line is as follows:</p>
<pre><code>**move test.txt C:\**
</code></pre>
<p>The above command moves the file <strong>test.txt</strong> to the root of the **C:** directory.</p>
<p>Arguments in command lines are related in the way that the co-occurrence of certain values can be indicative of malicious activity. NLP models are worth exploring here since these models are designed to understand and interpret relationships in natural (human) language, and since command lines often use some natural language.</p>
<h2 id="finetuningahuggingfacemodel">Fine-tuning a Hugging Face model</h2>
<p>Hugging Face is a data science platform that provides tools for machine learning (ML) enthusiasts to build, train, and deploy ML models using open source code and technologies. Its model hub has a wealth of models, trained for a variety of NLP tasks. You can either use these pre-trained models as-is to make predictions on your data, or fine-tune the models on datasets specific to your <a href="https://www.ibm.com/cloud/learn/natural-language-processing">NLP</a> tasks.</p>
<p>The first step in fine-tuning is to instantiate a model with the model configuration and pre-trained weights of a specific model. Random weights are assigned to any task-specific layers that might not be present in the base model. Once initialized, the model can be trained to learn the weights of the task-specific layers, thus fine-tuning it for your task. Hugging Face has a method called <a href="https://huggingface.co/docs/transformers/v4.21.1/en/main_classes/model#transformers.PreTrainedModel.from_pretrained">from_pretrained</a> that allows you to instantiate a model from a pre-trained model configuration.</p>
<p>For our command line classification model, we created a <a href="https://huggingface.co/docs/transformers/model_doc/roberta">RoBERTa</a> model instance with encoder weights copied from the <a href="https://huggingface.co/roberta-base">roberta-base</a> model, and a randomly initialized sequence classification head on top of the encoder:</p>
<p><strong>model = RobertaForSequenceClassification.from_pretrained('roberta-base', num_labels=2)</strong></p>
<p>Hugging Face comes equipped with a <a href="https://huggingface.co/docs/transformers/v4.21.0/en/main_classes/tokenizer">Tokenizers</a> library consisting of some of today's most used tokenizers. For our model, we used the <a href="https://huggingface.co/docs/transformers/model_doc/roberta#transformers.RobertaTokenizer">RobertaTokenizer</a> which uses <a href="https://en.wikipedia.org/wiki/Byte_pair_encoding">Byte Pair Encoding</a> (BPE) to create tokens. This tokenization scheme is well-suited for data belonging to a different domain (command lines) from that of the tokenization corpus (English text). A code snippet of how we tokenized our dataset using <strong>RobertaTokenizer</strong> can be found <a href="https://gist.github.com/ajosh0504/4560af91adb48212402300677cb65d4a#file-tokenize-py">here</a>. We then used Hugging Face's <a href="https://huggingface.co/docs/transformers/v4.21.0/en/main_classes/trainer#transformers.Trainer">Trainer</a> API to train the model, a code snippet of which can be found <a href="https://gist.github.com/ajosh0504/4560af91adb48212402300677cb65d4a#file-train-py">here</a>.</p>
<p>ML models do not understand raw text. Before using text data as inputs to a model, it needs to be converted into numbers. Tokenizers group large pieces of text into smaller semantically useful units, such as (but not limited to) words, characters, or subwords — called token —, which can, in turn, be converted into numbers using different encoding techniques.</p>
<blockquote>
  <ul>
  <li>Check out <a href="https://youtu.be/_BZearw7f0w">this</a> video (2:57 onwards) to review additional pre-processing steps that might be needed after tokenization based on your dataset.</li>
  <li>A complete tutorial on how to fine-tune pre-trained Hugging Face models can be found <a href="https://huggingface.co/docs/transformers/training">here</a>.</li>
  </ul>
</blockquote>
<h2 id="importingcustommodelsintoelastic">Importing custom models into Elastic</h2>
<p>Once you have a trained model that you are happy with, it's time to import it into Elastic. This is done using <a href="https://www.elastic.co/guide/en/elasticsearch/client/eland/current/machine-learning.html">Eland</a>, a Python client and toolkit for machine learning in Elasticsearch. A code snippet of how we imported our model into Elastic using Eland can be found <a href="https://gist.github.com/ajosh0504/4560af91adb48212402300677cb65d4a#file-import-py">here</a>.<br />
You can verify that the model has been imported successfully by navigating to <strong>Model Management \&gt; Trained Models</strong> via the Machine Learning UI in Kibana:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt989029e38ac28371/6a7d80dc4c4bfb9a5ecca843/Imported_model_in_the_Trained_Models_UI.png" alt="Imported model in the Trained Models UI" title="Imported model in the Trained Models UI" /></p>
<h2 id="usingthetransformermodelforinferenceaseriesofexperiments">Using the Transformer model for inference — a series of experiments</h2>
<p>We ran a series of experiments to evaluate whether or not our Transformer model could be used for real-time inference. For the experiments, we used a dataset consisting of ~66k command lines.</p>
<p>Our first inference run with our fine-tuned <strong>RoBERTa</strong> model took ~4 hours on the test dataset. At the outset, this is much slower than the tree-based model that we were trying to beat at ~3 minutes for the entire dataset. It was clear that we needed to improve the throughput and latency of the PyTorch model to make it suitable for real-time inference, so we performed several experiments:</p>
<h3 id="usingmultiplenodesandthreads">Using multiple nodes and threads</h3>
<p>The latency numbers above were observed when the models were running on a single thread on a single node. If you have multiple Machine Learning (ML) nodes associated with your Elastic deployment, you can run inference on multiple nodes, and also on multiple threads on each node. This can significantly improve the throughput and latency of your models.</p>
<p>You can change these parameters while starting the trained model deployment via the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/master/start-trained-model-deployment.html">API</a>:</p>
<pre><code>**POST \_ml/trained\_models/\\&lt;model\_id\\&gt;/deployment/\_start?number\_of\_allocations=2&amp;threa ds\_per\_allocation=4**
</code></pre>
<p><strong>number_of_allocations</strong> allows you to set the total number of allocations of a model across machine learning nodes and can be used to tune model throughput. <strong>threads_per_allocation</strong> allows you to set the number of threads used by each model allocation during inference and can be used to tune model latency. Refer to the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/master/start-trained-model-deployment.html">API documentation</a> for best practices around setting these parameters.</p>
<p>In our case, we set the <strong>number_of_allocations</strong> to <strong>2</strong> , as our cluster had two ML nodes and <strong>threads_per_allocation</strong> to <strong>4</strong> , as each node had four allocated processors.</p>
<p>Running inference using these settings <strong>resulted in a 2.7x speedup</strong> on the original inference time.</p>
<h3 id="dynamicquantization">Dynamic quantization</h3>
<p>Quantizing is one of the most effective ways of improving model compute cost, while also reducing model size. The idea here is to use a reduced precision integer representation for the weights and/or activations. While there are a number of ways to trade off model accuracy for increased throughput during model development, <a href="https://pytorch.org/tutorials/intermediate/dynamic_quantization_bert_tutorial.html">dynamic quantization</a> helps achieve a similar trade-off after the fact, thus saving on time and resources spent on iterating over the model training.</p>
<p>Eland provides a way to dynamically quantize your model before importing it into Elastic. To do this, simply pass in quantize=True as an argument while creating the TransformerModel object (refer to the code snippet for importing models) as follows:</p>
<pre><code>**# Load the custom model**
**tm = TransformerModel("model", "text\_classification", quantize=True)**
</code></pre>
<p>In the case of our command line classification model, we observed the model size drop from 499 MB to 242 MB upon dynamic quantization. Running inference on our test dataset using this model <strong>resulted in a 1.6x speedup</strong> on the original inference time, for a slight drop in model <a href="https://en.wikipedia.org/wiki/Sensitivity_and_specificity"><strong>sensitivity</strong></a> (exact numbers in the following section) <strong>.</strong></p>
<h3 id="knowledgedistillation">Knowledge Distillation</h3>
<p><a href="https://towardsdatascience.com/knowledge-distillation-simplified-dd4973dbc764">Knowledge Distillation</a> is a way to achieve model compression by transferring knowledge from a large (teacher) model to a smaller (student) one while maintaining validity. At a high level, this is done by using the outputs from the teacher model at every layer, to backpropagate error through the student model. This way, the student model learns to replicate the behavior of the teacher model. Model compression is achieved by reducing the number of parameters, which is directly related to the latency of the model.</p>
<p>To study the effect of knowledge distillation on the performance of our model, we fine-tuned a <a href="https://huggingface.co/distilroberta-base">distilroberta-base</a> model (following the same procedure described in the fine-tuning section) for our command line classification task and imported it into Elastic. <strong>distilroberta-base</strong> has 82 million parameters, compared to its teacher model, <strong>roberta-base</strong> , which has 125 million parameters. The model size of the fine-tuned <strong>DistilRoBERTa</strong> model turned out to be <strong>329</strong> MB, down from <strong>499</strong> MB for the <strong>RoBERTa</strong> model.</p>
<p>Upon running inference with this model, we <strong>observed a 1.5x speedup</strong> on the original inference time and slightly better model sensitivity (exact numbers in the following section) than the fine-tuned roberta-base model.</p>
<h3 id="dynamicquantizationandknowledgedistillation">Dynamic quantization and knowledge distillation</h3>
<p>We observed that dynamic quantization and model distillation both resulted in significant speedups on the original inference time. So, our final experiment involved running inference with a quantized version of the fine-tuned <strong>DistilRoBERTa</strong> model.</p>
<p>We found that this <strong>resulted in a 2.6x speedup</strong> on the original inference time, and slightly better model sensitivity (exact numbers in the following section). We also observed the model size drop from <strong>329</strong> MB to <strong>199</strong> MB after quantization.</p>
<h2 id="bringingitalltogether">Bringing it all together</h2>
<p>Based on our experiments, dynamic quantization and model distillation resulted in significant inference speedups. Combining these improvements with distributed and parallel computing, we were further able to <strong>reduce the total inference time on our test set from four hours to 35 minutes</strong>. However, even our fastest transformer model was still several magnitudes slower than the tree-based model, despite using significantly more CPU resources.</p>
<p>The Machine Learning team here at Elastic is introducing an inference caching mechanism in version 8.4 of the Elastic Stack, to save time spent on performing inference on repeat samples. These are a common occurrence in real-world environments, especially when it comes to Security. With this optimization in place, we are optimistic that we will be able to use transformer models alongside tree-based models in the future.</p>
<p>A comparison of the sensitivity (true positive rate) and specificity (true negative rate) of our tree-based and transformer models shows that an ensemble of the two could potentially result in a more performant model:</p>
<p>|                         |                 |                         |                 |                         |
| ----------------------- | --------------- | ----------------------- | --------------- | ----------------------- |
| Model                   | Sensitivity (%) | False Negative Rate (%) | Specificity (%) | False Positive Rate (%) |
| Tree-based              | 99.53           | 0.47                    | 99.99           | 0.01                    |
| RoBERTa                 | 99.57           | 0.43                    | 97.76           | 2.24                    |
| RoBERTa quantized       | 99.56           | 0.44                    | 97.64           | 2.36                    |
| DistilRoBERTa           | 99.68           | 0.32                    | 98.66           | 1.34                    |
| DistilRoBERTa quantized | 99.69           | 0.31                    | 98.71           | 1.29                    |</p>
<p>As seen above, the tree-based model is better suited for classifying benign data while the transformer model does better on malicious samples, so a weighted average or voting ensemble could work well to reduce the total error by averaging the predictions from both the models.</p>
<h2 id="whatsnext">What's next</h2>
<p>We plan to cover our findings from inference caching and model ensembling in a follow-up blog. Stay tuned!</p>
<p>In the meanwhile, we’d love to hear about models you're building for inference in Elastic. If you'd like to share what you're doing or run into any issues during the process, please reach out to us on our <a href="https://ela.st/slack">community Slack channel</a>and <a href="https://discuss.elastic.co/c/security">discussion forums</a>. Happy experimenting!</p>]]></content:encoded>
    <link>https://www.elastic.co/security-labs/blog/getting-the-most-out-of-transforms-in-elastic</link>
    <guid isPermaLink="false">getting-the-most-out-of-transforms-in-elastic</guid>
    <category><![CDATA[Security Operations]]></category>
    <dc:creator><![CDATA[Apoorva Joshi,Thomas Veasey,Benjamin Trent]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdba51d7bbafe2267/6a7d80dfc2cc09aa1a246682/machine-learning-1200x628px-2021-notext.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 23 Aug 2022 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>