Security
Software and Technology

1 billion events a day, 900+ detection rules as code: How a months-old team built Chainguard's security operations on Elastic Security

See story snapshot
  • ~1B
    Events a day ingested on Elastic Security
  • 900+
    Enabled detection rules, 280+ custom, managed as code
  • 32
    Log sources collecting; 27 live without custom code
  • 21 of 26
    Critical sources covered, remaining gaps tracked in inventory

Chainguard, the trusted source for hardened, production-ready open source, runs detection and response on Elastic Security Serverless, managing more than 900 detection rules as version-controlled code and enriching them with entity context, built by a cyber resiliency team that was only months old.

Summary

Chainguard is the trusted source for open source, delivering hardened, secure, production-ready builds of the software that organizations and AI agents depend on, with customers including Anduril, Canva, Fortinet, Hewlett Packard Enterprise, OpenAI, Snap Inc., and Snowflake. Its newly formed cyber resiliency team inherited a fast-moving, fragmented environment with no single place where an investigation lived end to end. The team standardized on Elastic Security running serverless, where it now ingests roughly a billion events a day, runs more than 900 detection rules managed as version-controlled code, and enriches them with entity context to catch threats like remote-worker fraud. Onboarding is mostly no-code, with 27 of its 32 sources arriving without custom collectors, and the team is testing an Elastic Agent Builder triage skill that reconstructs each case and recommends a verdict for a human to decide.

A new team standing up security at an engineering-driven company

Chainguard operates in an unusual position: a security company that other security-conscious organizations depend on, shipping hardened open source images to customers whose own users can include government agencies. How it runs its own security operations is a first-order concern, not a back-office function.

The cyber resiliency team responsible for detection and response is new; most of its members joined within the past year. They inherited an environment moving as fast as the company around it: dozens of log sources, hundreds of alerts a day, and no consolidated place where an investigation lived end to end. Institutional knowledge was thin because both the team and the environment were young.

"We get at least 100 alerts a day, and it takes a lot of time and effort to investigate every one. There are a lot of fast-moving pieces, and we're trying to keep up."

– Joaquin Nafstad, Senior Security Engineer, Chainguard

The goal was not simply to survive the alert volume. It was to build security operations a new engineer could pick up and extend without reverse-engineering someone's chat history, and to keep engineers on engineering rather than hand-triaging alerts. That meant standardizing on one platform first and then layering automation on top.

Choosing Elastic

Chainguard is cloud-native by design, with nothing on-premises, so any security platform had to be vendor-managed cloud infrastructure the team would not operate itself. Elastic Security Serverless fit that constraint, and the pitch to leadership was simple: The team had no capacity to manage clusters or tune performance.

The platform did not walk into an empty field. The team ran a structured evaluation, scoring the options it had in hand against the log sources it needed to cover and the criteria that mattered most: whether a platform could accommodate their critical log sources, ease of integration, strength of search and query language, the ability to build custom detections, and preventative and response automation. Across that matrix, Elastic came out strongest for the security-operations role.

What turned a month-to-month contract into a multiyear commitment was less any single feature than direction. Elastic's move toward AI-first security operations, GenAI-assisted detection, entity analytics, and agentic workflows reframed it in the team's eyes from a conventional SIEM into a strategic platform to build on.

"I built a spreadsheet with every log source we needed to cover down one axis and the solutions we had in hand across the other, then quantified it. I even shared it with Elastic when we were all on-site."

–  Mike Behrmann, Director of Cyber Resiliency, Chainguard

The partnership motion mattered, too. Elastic brought senior product, security, and engineering leaders into the room; ran hands-on workshops on detection development, telemetry ingestion, and normalization; and gave the team direct access to its own practitioners. For a new team that knew prior underutilization was partly a staffing and expertise gap, that enablement was what closed the gap between the platform's potential and what the team could actually run.

The team standardized on Elastic and moved fast. New to the platform, they leaned on Elastic's prebuilt detection rules for quick coverage, then invested in the deeper work of standardizing every source to a common schema and getting fluent in the query language, which they found markedly more powerful for building custom detections.

Inside Chainguard's Elastic deployment

Chainguard runs Elastic Security as the backbone of detection and response, normalizing every source to a common schema (ECS) on the way in. The end-to-end flow, at operational altitude:

  • Sources: 32 log sources collect today, across identity, endpoint, cloud, source control, and collaboration. Most arrive with no custom code: packaged Elastic integrations (several fully agentless) and configurable inputs cover 27 of the 32; only 5 need a custom collector for a vendor API without a native integration. Coverage spans 21 of the team's 26 critical sources, with the remaining gaps tracked in a formal inventory.
  • Detections: More than 900 detection rules run across dedicated spaces. Roughly two-thirds are Elastic's own prebuilt content from Elastic Security Labs, and the rest are custom, most authored in the Elasticsearch Query Language (ES|QL). Custom lookup indices and Elastic's Entity Analytics add entity context and enrichment to prioritize higher-risk activity, so rules can weigh identity and endpoint signals together rather than firing in isolation.
  • Cases: Elastic's detections feed the team's case-management layer, which runs roughly 150 to 200 cases a day. About 60% are resolved automatically within minutes, and analysts work the rest. Elastic remains the evidence store behind every case.
  • Triage: High-confidence benign cases are closed automatically. Cases below the team's confidence bar go to an analyst, supported by a triage skill built on Elastic Agent Builder that reconstructs each case and recommends a verdict. The skill is in testing (see below).
  • Escalation: Confirmed findings route out to the team's incident-management, ticketing, and chat tools for action.

Technical highlights

  • Elastic Security fully managed, every source normalized to ECS, no self-managed clusters
  • 32 sources collecting; 27 onboarded without custom code, 5 via custom collectors; 21 of 26 critical sources covered
  • 900+ detection rules (280+ custom, about two-thirds Elastic Security Labs content) managed as detection-as-code in git with schema and ES|QL validation and CI deploy per space
  • Entity Analytics and lookup indices add entity context and enrichment to prioritize higher-risk activity
  • Agent Builder triage skill, in testing, reads cases and recommends human-reviewed verdicts
  • Query latency at scale: searches typically return in under a second, with heavy aggregations across a full day taking one to two seconds

The capabilities

Detection-as-code: 900+ rules an engineer can actually read

Chainguard does not manage detections by clicking around a console. Its rules, more than 900 enabled and 280+ of them custom, live in a git repository and ship through a pipeline: a pull request runs schema and ES|QL validation and an API-contract check and verifies the query against live data with a read-only key, and then, on merge, CI imports the changed rules into the right space. Anyone can see how a detection works and why it fires without archaeology across chat threads and personal notes, and exceptions and tuning are themselves version-controlled changes rather than silent edits.

A detection they are proud of: Catching remote-worker fraud

One custom detection shows what the entity-enriched approach buys them. It correlates sign-ins across the team's identity and productivity platforms with known device-farm, proxy, and remote-KVM infrastructure, using entity context and enrichment to prioritize higher-risk activity, surfacing the kind of remote-IT-worker and laptop-farm fraud that has become a live threat for fast-hiring companies. The detection only works because identity and endpoint activity can be weighed against that entity context in one place.

An Agent Builder triage skill that shows its work

Beyond the prebuilt skills the team uses ad hoc for threat hunting, entity analytics, and anomaly detection, Chainguard has built its own triage skill on Elastic Agent Builder. High-confidence benign cases are closed automatically, and the automation sometimes reaches out to a user over chat to confirm activity. Cases that fall below the team's confidence bar go to an analyst, supported by the skill: invoked on a case, it reads the case and its underlying alert, the affected entities, the surrounding Elastic activity, and relevant context, then reconstructs what happened, grades the evidence, and recommends a standardized verdict, posting the supporting evidence back to the case under the analyst's name. The skill is in testing today. Notably, the same platform the team's engineers use to build these skills in Agent Builder is the one the security team runs detection and response on, one data model rather than a separate stack bolted alongside.

"We're a security company, so anything we do needs to be audited and documented. Anything our AI SOC isn't more than 95% confident about gets routed to a human for approval."

– Joaquin Nafstad, Senior Security Engineer, Chainguard

That discipline is the design, not an afterthought: deterministic queries and workflows for anything repeatable, large language model (LLM) reasoning reserved for synthesis such as connecting related signals or finding gaps, and a person on anything uncertain or potentially destructive. Keeping the case, the context, and the decision in one place is what makes the human review fast rather than a context-switch across tools.

Onboarding without the custom code

Getting a new source in is mostly a no-code exercise. Where a packaged integration exists and the source owner can supply credentials, the team can bring a source in and verify data is flowing in minutes; configurable inputs cover most of the rest, and only a handful of vendor APIs without a native integration need a custom collector.

"For an agentless integration, if we can get credentials to bring the log sources in, it takes less than 10 minutes. That is the biggest time saver ever."

– Joaquin Nafstad, Senior Security Engineer, Chainguard

Serverless: No clusters to manage

Running on serverless means the team spends no time on cluster sizing, performance tuning, or infrastructure operations, even at the team's ingest volume. For integrations the team hosts itself, heavy log volume occasionally slows its own collectors, but for anything Elastic manages, scaling is not something the team has to think about. Offloading platform operations entirely was the point.

Before Elastic, and with Elastic today


Before ElasticWith Elastic today
Log sourcesScattered across many vendors, no single home32 sources on one platform, 27 onboarded without custom code
Critical coverageGaps while the team ramped21 of 26 critical sources, remaining gaps tracked
DetectionsLimited, inconsistent900+ rules (280+ custom) managed as detection-as-code
Alert loadManual investigation of nearly every alert150 to 200 cases a day; about 60% resolved automatically within minutes, analysts work the rest
TriageEntirely by handHigh-confidence benign cases auto-close; the rest go to an analyst, supported by the Agent Builder skill (in testing)
InfrastructureWould have required managing clustersServerless, nothing for the team to operate

Getting to value fast

Two things stand out about how the team got this far this quickly. First, it invested early in the fundamentals, standardizing sources to a common schema and getting fluent in ES|QL, which paid off in how much it could then build. Second, it built its own mechanisms to keep pace with a platform that ships constantly, including a dedicated internal channel and an RSS subscription to Elastic's security releases, so new capabilities stopped slipping past unnoticed.

The team also drew directly on how Elastic runs security internally. Sessions with Elastic's own InfoSec team, showing how they use Workflows automation in production as customer zero, carried weight with Chainguard's leadership in a way that product demos do not.

"It helps a lot to see how a bigger company uses its own tool as customer zero. That was eye-opening for our leadership, seeing how an actual security team does security operations in production."

– Joaquin Nafstad, Senior Security Engineer, Chainguard

What comes next?

With detection-as-code running and entity-enriched detections in production, the near-term work is to take the triage skill from testing into everyday operation across the full case queue and to close the remaining critical-source gaps. The team is also evaluating Elastic's native Attack Discovery, where an alert analysis workflow first cuts alert noise and token usage, then Attack Discovery investigates what remains into attack narratives, with Workflows automating the response and a human keeping the decisions. That model-agnostic, cost-aware path, matching models to tasks using Elastic's published benchmarking, maps directly to Chainguard's stated direction and its focus on cost control.

The bigger arc is consolidation. The team piloted a data-loss automation on Elastic Workflows that ran end to end inside Elastic, then paused it when the alert volume drove too much model usage; the lesson, to cut noise deterministically before sending alerts to a model, now shapes how the team is approaching native triage. It also wants more of its institutional knowledge captured in indexes that agents can draw on in future investigations, the shareable, repeatable operating model it is building toward. For now, the plan is to migrate the most straightforward automation first, prove it against a review threshold, and consolidate onto Elastic over roughly the next six months rather than in a single cutover.

Notably, Chainguard is not rebuilding a traditional tiered SOC. Rather than a base of analysts triaging alerts and escalating upward, the team is designing toward prioritized investigations, deterministic workflows for repeatable situations, and specialized agents that loop in a human only when confidence is low.

Your organization may not be standing up security operations at Chainguard's pace, but the same principles apply whether you are onboarding your first critical source or running a billion events a day across a cloud-native environment.

Topics: Agentic security operations, Elastic Security, Agent Builder, Workflows, Entity Analytics, detection-as-code, ES|QL, ECS, insider risk, serverless, bring-your-own-model, SIEM, Software & Technology

See how Elastic Security brings detection, triage, and response into one agentic security operations platform, or start now with a free trial.