Large language model performance matrix for Elastic Security

This page summarizes internal test results comparing large language models (LLMs) across Elastic Security AI chat and AI-powered feature use cases. The matrix tests each model across Agent Builder, Attack Discovery, and Automatic Migration. To learn more about these use cases, refer to AI-powered features. To learn how these scores are produced, refer to Benchmarking the Agentic SOC on Elastic Security Labs.

Important

Higher scores indicate better performance, on a scale of 1 to 10. A score of 10 on a capability means the model met or exceeded all task-specific benchmarks for that capability.

Any model that scores 5 or below for a capability is not recommended for that task.

The matrix uses three top-line capability scores — Agent Builder, Attack Discovery, and Automatic Migration — that roll up into a single Overall Score. You can read the table top-down, from "how does this model perform across our AI features?" to "how good is it at the specific job I care about?"

  • Overall Agent Builder Score is the average of the seven Agent Builder sub-capabilities. It summarizes how well a model handles agentic Security work end to end.
  • Overall Score is the average of the Agent Builder, Attack Discovery, and Automatic Migration scores. It reflects how a model performs across the breadth of Elastic's AI features rather than any single workflow, and is the default sort for the tables below.

For a full walkthrough of the evaluation framework behind these scores — the synthetic intrusion, deterministic ground truth, trace-level grading, and blind judging — refer to Benchmarking the Agentic SOC on Elastic Security Labs.

  • Alert Analysis — Triage an alert, reach the correct disposition, pull related alerts, and enrich with threat intel.
  • Entity Analytics — Investigate hosts and users using purpose-built entity lookups and risk context.
  • Threat Hunting — Generate and run queries against process, file, and network telemetry to find specific hunt artifacts.
  • Detection Rules — Author a working detection rule, grounded in research where requested.
  • Workflow Authoring — Produce a valid, executable automation workflow (verified by actually creating, enabling, and running it).
  • Triggering Workflows — Call the correct backed action for the task (for example, a hash lookup, an on-call schedule, or case creation).
  • Multi-Step Executions — Chain several steps in the right order, carrying findings forward, without skipping or fabricating steps.

Models from third-party LLM providers.

Model Agent Builder: Alert Analysis Agent Builder: Entity Analytics Agent Builder: Threat Hunting Agent Builder: Detection Rules Agent Builder: Workflow Authoring Agent Builder: Triggering Workflows Agent Builder: Multi-Step Executions Overall Agent Builder Score Attack Discovery Automatic Migration Overall Score
Anthropic Claude Sonnet 4.5 9.00 8.00 7.00 8.00 4.00 9.00 8.00 7.57 9.40 9.61 8.86
Anthropic Claude Opus 4.7 5.00 8.00 8.00 8.00 6.00 9.00 8.00 7.43 9.10 9.90 8.81
Anthropic Claude Opus 4.6 9.00 8.00 8.00 8.00 4.00 8.00 8.00 7.57 9.20 9.61 8.79
Anthropic Claude Sonnet 4.6 9.00 8.00 8.00 8.00 6.00 9.00 8.00 8.00 9.00 9.23 8.74
Anthropic Claude Opus 5 6.00 8.00 8.00 8.00 9.00 9.00 6.00 7.71 9.60 8.75 8.69
Anthropic Claude Opus 4.8 7.00 7.00 8.00 8.00 7.00 9.00 6.00 7.43 8.00 10.00 8.48
Anthropic Claude Sonnet 5 8.00 8.00 7.00 8.00 7.00 9.00 5.00 7.43 9.00 8.84 8.42
OpenAI GPT-5.2 6.00 7.00 8.00 8.00 6.00 8.00 8.00 7.29 8.00 9.60 8.30
OpenAI GPT-5.5 8.00 8.00 7.00 8.00 9.00 8.00 8.00 8.00 8.20 8.55 8.25
Anthropic Claude Opus 4.5 8.00 8.00 8.00 8.00 6.00 9.00 8.00 7.86 8.70 8.17 8.24
Google Gemini 2.5 Pro 5.00 5.00 7.00 6.00 6.00 9.00 8.00 6.57 8.70 9.32 8.20
OpenAI GPT-5.6 Terra 7.00 7.00 7.00 8.00 5.00 8.00 8.00 7.14 6.50 9.61 7.75
OpenAI GPT-5.4 7.00 7.00 8.00 7.00 7.00 9.00 8.00 7.57 5.30 9.84 7.57
Google Gemini 3.6 Flash 8.00 5.00 7.00 8.00 7.00 8.00 8.00 7.29 8.20 7.21 7.57
OpenAI GPT-5.6 Sol 6.00 8.00 7.00 8.00 9.00 8.00 6.00 7.43 8.00 7.21 7.55
Anthropic Claude Haiku 4.5 3.00 8.00 7.00 6.00 7.00 9.00 8.00 6.86 7.00 8.65 7.50
Google Gemini 3.0 Flash 7.00 6.00 7.00 8.00 3.00 8.00 8.00 6.71 6.00 9.71 7.47
OpenAI GPT-5.6 Luna 8.00 7.00 6.00 8.00 9.00 8.00 8.00 7.71 6.30 7.88 7.30
Google Gemini 2.5 Flash 3.00 4.00 5.00 3.00 3.00 6.00 6.00 4.29 6.80 9.61 6.90
Google Gemini 3.5 Flash 8.00 6.00 7.00 7.00 9.00 8.00 8.00 7.57 6.30 6.73 6.87
Google Gemini 3.1 Flash Lite 8.00 7.00 7.00 8.00 9.00 8.00 6.00 7.57 2.50 9.51 6.53
OpenAI GPT-5.4 Nano 4.00 4.00 5.00 4.00 9.00 7.00 6.00 5.57 3.50 8.75 5.94
Google Gemini 3.1 Pro (Preview) 8.00 5.00 7.00 8.00 7.00 9.00 7.00 7.29 3.20 6.44 5.64
OpenAI GPT-5.4 Mini 4.00 6.00 6.00 5.00 6.00 9.00 8.00 6.29 1.00 9.51 5.60
Google Gemini 3.5 Flash Lite 5.00 3.00 6.00 7.00 7.00 8.00 7.00 6.14 0.00 9.81 5.32
Google Gemini 2.5 Flash Lite 5.00 2.00 1.00 2.00 0.00 8.00 5.00 3.29 0.00 6.15 3.15

Models you can deploy yourself, ranked by overall score.

Last measured in September 2026.

Model Agent Builder: Alert Analysis Agent Builder: Entity Analytics Agent Builder: Threat Hunting Agent Builder: Detection Rules Agent Builder: Workflow Authoring Agent Builder: Triggering Workflows Agent Builder: Multi-Step Executions Overall Agent Builder Score Attack Discovery Automatic Migration Overall Score
Z.ai GLM 5.3 8.00 7.00 7.00 8.00 9.00 9.00 8.00 8.00 8.00 7.30 7.77
Gemma 4 31B IT 6.00 7.00 7.00 8.00 3.00 9.00 8.00 6.86 4.30 9.61 6.92
Kimi K2.6 8.00 6.00 7.00 8.00 9.00 9.00 8.00 7.86 8.00 2.31 6.05
DeepSeek V4 Pro 6.00 8.00 7.00 8.00 9.00 7.00 8.00 7.57 3.50 6.25 5.77
Qwen 3.8 2.4T A95B 8.00 7.00 7.00 9.00 9.00 9.00 9.00 8.29 0.00 7.88 5.39
OpenAI GPT-OSS 120B 1.00 1.00 1.00 3.00 5.00 8.00 1.00 2.86 2.00 9.51 4.79
Qwen 3.6 27B 6.00 7.00 8.00 7.00 7.00 9.00 8.00 7.43 0.00 5.67 4.37
OpenAI GPT-OSS 20B 2.00 2.00 2.00 5.00 4.00 7.00 5.00 3.86 1.00 6.05 3.64