Threat Command

Behind the tags: How Elastic SIEM grades 1,781 detection rules on noise, speed, and threat coverage

This article explains how Elastic SIEM uses a monthly automated telemetry pipeline to score prebuilt detection rules across noise, performance, threat, and profile dimensions, helping security teams decide which rules to enable first.

Of Elastic's 2100+ prebuilt detection rules, 385 are tagged Profile: Recommended and 296 are tagged Aggressive. The rest (62% of the list) are deliberately left untagged. Prebuilt rule now carries tags like Noise: Low, Performance: Fast, and Profile: Recommended, recalculated every month from real fleet telemetry rather than fixed once and left alone. This post covers how the scoring actually works and how to use the tags to decide what to enable first.

Elastic SIEM rule metadata has two jobs

Most of the prebuilt rule metadata answers the question What does this rule detect? The metadata includes the MITRE ATT&CK technique and the data source, along with the platform. Together, these provide static context.

The new Noise, Performance, Threat, and Profile tags answer something different: How does this rule behave in production? We now have the answer to this question that users have been asking. This operational context changes over time as environments grow, adoption shifts, and detection patterns evolve. A rule that's quiet today might become noisy after a platform update. A rule tagged Noise: Unknown when it first ships will eventually accumulate enough fleet signal to get a real label.

Updated monthly using real-world usage data, these tags stay fresh and deliver actionable insights, and that’s why they’re so valuable.

How Elastic SIEM's monthly tag refresh works

We love data and automation. Once a month, an automated pipeline reads a 30-day window of telemetry across Elastic deployments and scores every prebuilt rule across four dimensions: Noise, Performance, Threat, and Profile.

The inputs are:

  • Analytics telemetry. Alert volume, distinct cluster counts, and firing density per rule.

  • Execution metrics. Average, 95th-percentile, and maximum execution times.

  • The rule's own TOML. Severity, age, platform, and existing tags.

The scoring is entirely deterministic. We do use large language models (LLMs) in one place, but only as a fallback for Threat classification on rules where our heuristic catalog has nothing to say. We cover more on that below.

Every month, the pipeline automatically opens a draft pull request (PR) on elastic/detection-rules with the proposed updates. Our Elastic Security Labs Threat Command team reviews that PR before anything merges: which rules moved between Recommended and Aggressive, which changed Noise or Performance buckets, and which gained or lost Threat tags. The pipeline generates the proposal, and the team makes the call.

Noise: How Elastic SIEM measures alert volume across the fleet

We make a decision on Noise tag based on the total alert volume that a rule generates across the fleet over 30 days. The core metric we use internally is global_noise, which is roughly the mean number of alerts per distinct environment on days that the rule actually fired. This smooths out the difference between a rule that hammers one customer and one that fires occasionally across many.

We classify rules into four buckets:

  1. Noise: Low. Around the 25th percentile for its platform, or fewer than 100 total alerts in the window, or six-months-or-older with essentially no signal (we treat long-lived quiet rules as low noise rather than broken).

  2. Noise: High. At or above the 85th percentile, or extreme alert density relative to the number of distinct environments, or a new rule that came out loud in its first 60 days.

  3. Noise: Medium. Everything between Low and High.

  4. Noise: Unknown. Less than 60 days old with too little signal to tell "quiet" from "not yet widely enabled."

That last bucket is worth a bit more explanation. A new rule that barely generates any alerts isn't necessarily low noise; it might just not be enabled in many places yet. We deliberately hold back the Noise: Low label for the first 60 days to avoid misleading you.

For platforms with enough rules in the fleet, such as Windows, Linux, macOS, AWS, Azure, Google Cloud Platform (GCP), Google Workspace, Microsoft 365, Okta, and GitHub, we calculate percentile cuts per platform. Everything else uses a shared fleet fallback, because what counts as "noisy" is genuinely different across environments.

Performance: How expensive is this rule to run?

The Performance tag tracks how long a rule takes to run across different clusters, measured in milliseconds. We use absolute thresholds rather than percentile-based buckets because execution times tend to cluster tightly enough that relative buckets would mislead.

  • Performance: Very Slow. Average of five seconds or more, or p95 at 15 seconds or more, or 25+ environments with at least one 30-second-plus execution.

  • Performance: Slow. Average of two seconds or more, or p95 at five seconds or more, or 10+ environments with a 30-second-plus execution.

  • Performance: Fast. Average of 400ms or less and p95 under one second and no 30-second executions and at least five environments sampled.

  • Performance: Normal. Enough data to classify, but not Fast, Slow, or Very Slow.

  • Performance: Unknown. Fewer than 50 total executions or fewer than three sampled environments.

We evaluate in this order: Unknown first (no judgment on insufficient data) and then Very Slow, Slow, Fast, and Normal as the default for everything in the middle. Fast is intentionally strict; it requires clean p95 numbers and enough sampling to be confident, not just a good average.

Threat: Deterministic first, AI-assisted as a fallback

The Threat tag communicates the detection priorities defined by our Threat Research and Detection Engineering (TRaDE) team. These priorities are informed by confirmed true positives observed in our own telemetry, as well as attack techniques and behaviors trending publicly. Examples include Supply Chain attacks, ClickFix, Phishing Kits, remote monitoring and management (RMM) abuse, and similar emerging or high-priority threats. This gives readers a clear view into the threat areas where our detection engineering efforts are currently focused.

We maintain a managed catalog of these Threat categories and use heuristics to match rules to them. Where the heuristic finds a clear match, the tag is applied deterministically. Where it doesn't, we can optionally use LLMs to review the rule and suggest a Threat category, gated by human reviews. The model receives a compact description of the rule and returns candidates, which we then filter against our managed catalog and cap at five tags per rule. The model cannot invent categories outside the approved list.

This keeps AI in a supporting role. Noise, Performance, and Profile are all scored by code. AI helps fill classification gaps for Threat tags, where identifying the appropriate category often depends more on understanding the rule's behavior and context than on threshold math.

Another important detail is that we don't overwrite manually curated tags. If a rule already has Threat: Cobalt Strike or another malware-family label, the pipeline leaves it alone. The same applies to Profile: Beta and other Profile tags set manually. The automated refresh only updates the categories it owns.

Figure 1. Example filtering for Supply Chain rules using the Threat tag.

Profile: Elastic SIEM's summary score

The Profile tag is where everything comes together. It synthesizes severity, noise, performance, and threat coverage into a single deployment posture recommendation.

Factor

Points

Critical severity

+3

High severity

+2

Medium severity

+1

Low severity

0

Noise: Low

+3

Noise: Medium

+1

Noise: High

−2

Noise: Unknown

0

Performance: Fast

+2

Performance: Normal

+1

Performance: Slow

−1

Performance: Very Slow

−2

Performance: Unknown

0

Covers a managed Threat tag

+4

Profile: Recommended requires a score of 7 or higher, with two hard gates on top of that: any rule with Noise: High becomes Aggressive, regardless of score, and low-severity rules are never Recommended, regardless of score.

Threat coverage carries the highest weight (+4) because the tags in our managed catalog represent mechanics that we actively track and prioritize. A medium-severity rule with no Threat match typically can't reach a score of 7 without either low noise or fast performance to make up the gap. A high-severity rule with low noise and fast performance hits exactly 7 without any Threat tag (2 + 3 + 2).

Profile: Aggressive applies when a rule is High noise (the hard gate, regardless of score) or when the score falls below 2, except for low-severity rules, which get no Profile tag at all.

If a rule doesn't quite fit Recommended or Aggressive, we simply leave it without a Profile tag. This isn't a gap in our classification; it's intentional. Those rules are worth evaluating on their own terms, and we didn't want to force them into a bucket that oversimplifies the tradeoff.

Currently, 385 of our 1,781 out-of-the-box detection rules (21.6%) are tagged Profile: Recommended, while 296 (16.6%) are tagged Profile: Aggressive. The remaining 1,100 rules (61.8%) are intentionally untagged, reflecting rules that don't currently meet the criteria for either profile. That 1,781 total covers only rules that we can tune through Elastic SIEM rule artifacts. It excludes building block rules, machine learning detections, deprecated rules, promotion rules, threat intelligence rules, and the guided-onboarding test rule. The snapshot below shows how those profiles are distributed across severity and platform, along with the Threat categories currently represented in the ruleset:

Figure 2. Profile distribution across Elastic's current out-of-the-box detection rules. These numbers are a point-in-time snapshot and will evolve as rules, telemetry, and tracked threats change.

Which Elastic SIEM rules should I enable first?

If you're setting up detection coverage for the first time, start with Profile: Recommended. 

  • Profile: Recommended. These are the rules that we believe are safe to enable broadly: high enough signal quality and low enough operational cost that you won't immediately be buried in noise or slow down your cluster.

  • Profile: Aggressive.  Check the individual Noise: and Performance: tags before enabling. A rule that's Noise: High but Performance: Fast and maps to a tracked Threat might still be worth enabling with a suppression set. The profile score tells you whether a rule is broadly safe to enable; the individual tags tell you which specific tradeoff you're looking at.

  • Mid-band rules (no Profile tag). Don't ignore these. Some of our most targeted detections sit there, but they require more environment-specific evaluation before a broad rollout.

  • Noise: Unknown. Treat this as a watch item, not a veto. New rules need time to accumulate fleet signals. Give it a release cycle or two (two to four weeks).

For all of our existing uses, you’ll also notice other tags added to further align toward a predictable and standardized Elastic XDR/Elastic SIEM taxonomy for prebuilt tags. This taxonomy includes tags for: 

  • Domain. - The visibility plane the behavior lives in.

  • Platform. The ecosystem that the rule applies to.

  • Data Source. The log stream or integration.

  • Service. Optional specific component.

  • Vuln. Vulnerabilities and Common Vulnerabilities and Exposures (CVEs).

  • Resources. Investigation guides and LLM consumption, among others.

In the coming release, we’ll also include MITRE ATLAS tags to accompany MITRE ATT&CK.

With additional ways to splice your rule search categories, near future search features will allow for more granular questions about coverage, such as Show me recommended (Profile) rules for identity (Domain) coverage related to ConsentFix (Threat) for Entra ID Sign-In Logs (Data Source). Check out our docs, for a catalog of tags now available. In the futuree, we’ll begin removing old tags that no longer align to the taxonomy, reducing duplicate tags and providing a better search experience. 

How to give feedback on Elastic SIEM rule tags

Telemetry is voluntarily shared by our users, and it's what makes this whole system work. If you've turned on telemetry sharing, you're directly contributing to more accurate tags for everyone.

If you have feedback on how we're doing this, including rules that you think are miscategorized, thresholds that don't match your reality, or edge cases that we haven't thought about, reach out via our Slack Community or the Discuss forum. The scoring model is something we'll keep iterating on, and your experience in production is exactly the kind of input we want.

Check the Elastic Security capabilities on your deployment, or start your free trial.

Related Content

Cloud Threat Emulation on Autopilot: Context is Everything

Cloud Threat Emulation on Autopilot: Context is Everything

Terrance DeJesus
Linux Detection Engineering - Local Privilege Escalation

Linux Detection Engineering - Local Privilege Escalation

Ruben Groenewoud
How to correlate Kubernetes audit logs with container runtime data

How to correlate Kubernetes audit logs with container runtime data

Isai Anthony
Linux Detection Engineering - Fileless Execution

Linux Detection Engineering - Fileless Execution

Ruben Groenewoud

Elastic Security Labs Newsletter