Security
Digital Native

How a global AI planning platform runs security operations on Elastic at 2–3 TB a day

  • Unlimited
    Threat-hunting retention through searchable snapshots and the frozen tier, versus the 10–20 days storage-bound tools allow
  • 2–3 TB/day
    Security and observability telemetry ingested across Google Cloud, Azure, and AWS
  • 1 language
    SOC analysts generate detection queries with the AI Agent instead of queuing searches to engineering

A leading AI planning platform consolidated its security operations on Elastic Security, giving its SOC unlimited threat-hunting retention across 2–3 TB of daily telemetry pulled from Google Cloud, Azure, and AWS.

Summary

For the Fortune 500 companies that run their supply chains, commercial operations, and profit-and-loss decisions on this platform, what's on the line is their most sensitive data — as well as the trust that comes with handing it over. Before standardising on Elastic, the security team at this global AI planning platform wrote new detections by hand through an overloaded engineering queue, and retention capped at roughly 10–20 days made it hard to hunt across the long behavior windows that reveal an anomaly. Today, the company ingests 2–3 TB of telemetry a day across Google Cloud, Azure, and AWS into one platform. Analysts author detections in natural language through the Elastic AI Agent, while searchable snapshots let them hunt across effectively unlimited history held in the frozen tier, backed by cheap cloud object storage. Here's how they do it.

Protecting Fortune 500 data at cloud scale

This company builds the platform its clients run their businesses on — supply chain planning, demand planning, and demand forecasting for some of the largest enterprises in the world. When a Fortune 500 company hands over its planning data, it hands over what it buys, what it sells, and where its margins live, along with the compliance and regulatory obligations that come attached.

The company is cloud-native and growing fast, so its security has to move at the same speed and clear the bar the world's largest enterprises set. To do this, the security team needs to collect telemetry from wherever it lives, retain enough of it to investigate properly, and let analysts act on it without waiting on anyone.

"Most of the Fortune 500, and even Fortune 10, run their planning on us. My job is to protect their data, because their reputation, their money, and their compliance are all at stake."

– Senior Director of Security Operations

What a cloud-native SecOps team needed that storage-bound tooling couldn't give

This company wanted a foundation as modern as the rest of its architecture, not a set of point tools bolted together. Two requirements drove the decision. The first was collection. As a cloud-native business running across Google Cloud, Azure, and AWS, the team needed to pull telemetry from all three without custom engineering for each. Elastic is cloud-agnostic, so the same platform collects seamlessly from every provider.

The second was history. Detecting a behavior anomaly, whether in a user or across the network, means looking across a long window of activity, not just the past couple of weeks. Existing approaches tied retention directly to how much storage the team was willing to pay for, which in practice capped useful history at roughly 10–20 days. In practice, this means that the cost of full context comes due exactly when an investigation needs it most. The team needed a platform that broke the link between how much data it kept and how much it could search.

A storage ceiling and an engineering bottleneck

Running a cloud-native platform across three cloud providers, applications, and network devices generates logs, metrics, and events in many formats and locations. This caused two burdens for the SecOps team of this major AI planning company.

First, authoring detections was manual and centralized. When the SOC needed a new search, the request went to an engineering team that was already heavily loaded, so analysts depended on someone else to turn an idea into a working query.

Second, retention was bound to storage. Because keeping data meant paying for hot storage to hold it, the window needed to see how a user or the network actually behaves over time was too short. So, threat hunting was constrained not by the analysts' skill, but by how far back they could look.

Inside the platform: Cloud-agnostic ingest and searchable snapshots in the frozen tier

Routinely, the company ingests 2–3 TB of telemetry a day. Collection is cloud-agnostic: Elastic pulls logs, metrics, and events from Google Cloud, Azure, and AWS seamlessly, along with the applications and network devices running across that environment, and lands them in one platform used for both security operations and observability.

The load-bearing architectural choice is the frozen tier. Searchable snapshots keep the company’s full history in low-cost cloud object storage rather than on local disk, and the frozen node searches it on demand. Any cloud bucket integrates with the frozen tier, so the searchable window is set by the object store, not by the size of a local disk. At 2–3 TB a day, a 10, 20, or even 100 TB disk would only ever cover a handful of days. The frozen tier removes that ceiling.

Technical highlights

  • Cloud-agnostic collection from Google Cloud, Azure, and AWS into one platform
  • 2–3 TB of security and observability telemetry ingested per day
  • Searchable snapshots hold data in the frozen tier for effectively unlimited retention
  • Any cloud object-storage bucket integrates directly with the frozen tier
  • Searchable data lives in cloud object storage, so how much history is searchable isn't capped by local disk
  • Elastic AI Agent generates detection queries from natural language descriptions
  • Attack Discovery to identify, protect, detect, respond, recover, and govern coverage
  • Single consolidated platform for SIEM, threat hunting, incident management, and observability

The platform in practice

Cloud-agnostic telemetry ingest across 3 clouds

The company’s engineers work with data arriving from everywhere at once, and none of it is useful until it can be seen together. Elastic collects from Google Cloud, Azure, and AWS without provider-specific engineering, so logs, metrics, and events from across the cloud-native environment land in one place the team can investigate from. Keeping full-fidelity telemetry, rather than dropping it to control cost, is what lets an investigation start from a complete picture instead of a sampled one.

Threat hunting across unlimited history

This is where the frozen tier pays off. To spot a behavior anomaly, an analyst has to look across a long window, not 10 or 20 days, and searchable snapshots make that practical. The team hunts across effectively unlimited data held in the frozen tier and backed by object storage, with no rehydration wait to reach older data during an investigation. Search performance stays real-time even as the searchable window grows, because compute is no longer tied to how much history is on disk.

"To spot a behavior anomaly you have to look across a long period, not 10 or 20 days. Searchable snapshots let me hunt across unlimited data in the frozen node, and I can point it at any cloud bucket."

– Senior Director of Security Operations

Attack Discovery and proactive response

The team's goal is to get ahead of threats rather than react to them after the fact. Elastic's Attack Discovery helps the team work across the full response lifecycle, from identifying and protecting through detecting, responding to, and recovering from threats, with governance alongside. Holding the long history the frozen tier provides is what makes that proactive posture possible — so the signal that distinguishes normal from anomalous behavior lives in weeks and months of data, not days.

From noise to a detection, in plain English

There is a version of a security team's day that is mostly maintenance, like making sure data flows and tending the tooling instead of the threats.

Thankfully, it’s the kind of day this company’s analysts no longer have.

Before, a new detection meant a ticket to an overloaded engineering team and a wait. Now an analyst describes the threat to the AI Agent in plain English and gets back the exact query needed, and then reads and runs it. The analyst stays in control of the judgment while the platform removes the drudgery, and because everything lands in one view, an analyst who sees an alert already has the context around it rather than opening five tabs to assemble it.

"Before the AI assistant, writing new searches was manual and my engineering team was buried in it. Now an analyst just asks in plain English, and it gives back the exact query they need."

– Senior Director of Security Operations

Before and after


BeforeAfter
Detection authoringNew searches written by hand through an overloaded engineering queueAnalysts generate detections in plain English with the AI Agent and verify the exact query returned
Retention and threat huntingRoughly 10–20 days, capped by storage sizingEffectively unlimited history in the frozen tier via searchable snapshots, backed by cloud object storage
Data collectionTelemetry spread across clouds, applications, and network devicesCloud-agnostic ingest from Google Cloud, Azure, and AWS into one platform
Investigation flowContext assembled by hand across separate toolsOne unified view, so context is present when an alert appears
Analyst roleWaiting on engineering, tending tooling, hunting only across recent dataAuthoring detections directly, hunting across full history, focused on judgment and response

 

Built to scale with the business, not to be rebuilt

The company keeps winning bigger clients and pushing deeper into AI-driven analytics, and every new enterprise it takes on is more data to protect and more trust to keep.

The platform absorbs that growth without forcing the retention-versus-cost tradeoff that limited the team before: At 2–3 TB a day and rising, the searchable window is set by object storage, not by disk. As a cloud-native software company whose reputation rests on being a safe place to put planning data, its security and observability foundation now scales with the business instead of becoming the next thing that has to be torn out and redone.

Your organization may not be securing Fortune 500 supply chain data at 2–3 TB a day today, but the same principles apply whether you are bringing together your first handful of data sources or hunting across effectively unlimited history: Keep full context, and keep it searchable.

"It's one consolidated platform I can rely on, and when something breaks, their support is there around the clock to help me build the RCA and answer to my leadership."

– Senior Director of Security Operations

See how Elastic Security lets you hunt threats across your full history instead of the last few weeks of storage, or start now with a free trial.