<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Luca Wintergerst - Elastic Observability Labs]]></title>
    <description><![CDATA[Trusted security news & research from the team at Elastic.]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Luca Wintergerst - Elastic Observability Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltad972c1c27dbefc6/6a88d9782904ea5e8511d473/observability-labs-thumbnail.png</url>
      <link>https://www.elastic.co/observability-labs/author/luca-wintergerst</link>
    </image>
    <link>https://www.elastic.co/observability-labs/author/luca-wintergerst</link>
    <atom:link href="https://www.elastic.co/observability-labs/rss/author/luca-wintergerst.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Sun, 11 Oct 2026 05:42:19 GMT</lastBuildDate>
  <item>
    <title><![CDATA[Meet Elastic® nightshift AI SRE, your always-on SRE coworker]]></title>
    <description><![CDATA[Elastic's SRE team has been running Elastic nightshift against our own production environment. It picks up incidents that never triggered an alert and shows its working for every conclusion it reaches.]]></description>
    <content:encoded><![CDATA[<p>Today we're announcing the Private Preview of the Elastic® nightshift AI SRE, built into Elastic Observability to find, investigate, and fix production problems alongside your team. The Private Preview on Serverless is free during the promotional period.  <a href="https://www.elastic.co/observability/ai-sre/signup">Sign up for the waitlist here →</a></p><p>Production is outrunning the tools that watch it. Teams ship faster than ever, and more of that code is written by AI. Yet keeping production healthy is still manual work. Someone has to write the alert rules that catch problems and work out why each alert fired, and then they have to apply the fix. Much of what they need to know lives only in the minds of a few engineers. Every release changes what needs watching, and this work has become the bottleneck.</p><p>Elastic nightshift AI SRE works alongside your team, with the full context about your systems. It watches your telemetry around the clock and detects problems, including the unknown ones that no alert rule covers. It investigates incidents the way that an experienced site reliability engineer (SRE) would, testing hypotheses against your telemetry data, such as logs, metrics, traces, and profiles, ruling out dead ends and showing you evidence for the root cause. Then it recommends a fix to remediate the problem and applies it automatically, if you want it to.</p><p>At every step, it uses the most relevant contextabout your data and systems, built from your telemetry, code, runbooks, and more. It learns proactively from your data and prior incidents, along with enterprise knowledge. And, like you and your team, it investigates more accurately with every incident it detects or any alert it investigates. Plus, it can work with you in the Kibana UI, a terminal, an agent of your choice, or anywhere else you work. </p><p>When we were building this system, it was incredibly important that we earn your trust. If you've been using large language models (LLMs) for a while, you know how overly confident some tend to be. With Elastic nightshift AI SRE, instead of just getting the conclusion, you can follow all the steps that the agent took to get there and even drill down into the raw data, if you need to, as the following animation shows. 

You can check out the interactive version of this article <a href="https://elastic.co/blog/introducing-elastic-nightshift-ai-sre-interactive-post">here</a>.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta15e1887802c8a69/6ac738c5d056022ac8294ebc/image3.gif" alt="Animation of the Elastic nightshift AI SRE investigating an incident and showing each step of its reasoning" /><h2>How Elastic nightshift AI SRE works</h2><p>Elastic nightshift consists of four engines that work together (or you can use them on their own): Context, Detection, Investigation, and Remediation. For example, you can have it investigate alerts from the monitoring tools that you already use without turning on detection.</p><p><strong>Engine</strong></p><p><strong>What it does</strong></p><p><strong>What triggers it</strong></p><p><strong>What you get</strong></p><p>Context</p><p>Builds knowledge of your environment from telemetry, code, runbooks, and wikis</p><p>Runs continuously from the moment it has access to your data and systems</p><p>Entities, relationships, and a long-term memory of past incidents that the other three engines draw on</p><p>Detection</p><p>Evaluates telemetry continuously and flags what needs attention, with no alert rules to write or maintain</p><p>Always on when configured for data in Serverless</p><p>A significant event grouping related signals across services, with the parts of your system that it affects</p><p>Investigation</p><p>Tests hypotheses in parallel against logs, metrics, traces, and profiles, separating symptoms from causes</p><p>Any alert, fired in Elastic or a connected tool; optionally a significant event</p><p>A root cause in plain language, affected services, proposed actions, and the blind spots that it couldn’t verify</p><p>Remediation</p><p>Picks a workflow or connector action and fills in the parameters for this incident, or gives you commands to run</p><p>A completed investigation</p><p>A fix that you review before it runs (or apply automatically, if you choose)</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt797d70fcacc7319b/6ac738e2ea4fe9e7e60f3f08/image8.png" alt="Architecture of the Elastic nightshift AI SRE: context, detection, investigation and remediation engines" /><p>Most importantly, you don't have to move your data to use it. The agents run on Elastic Observability Serverless and reach out to your telemetry and alerts in the systems where they already live. By technical preview later this year, it will work with:</p><ul><li><p>Data in Elastic Observability Serverless (sending data there is optional).</p></li><li><p>Elastic Cloud Hosted (ECH) deployments.</p></li><li><p>Self managed clusters (provided that the Elasticsearch endpoint is exposed appropriately).</p></li><li><p>Popular third-party observability tools.</p></li><li><p>Popular telemetry stores and databases.</p></li></ul><p>When you sign up for the waitlist, tell us which third-party tools, telemetry stores, and databases you want most so we can prioritize accordingly. In addition, we’re working toward a plan to deploy all this functionality in self-managed environments. We’ll have more to share later this year.</p><h3>Context Engine: Giving an AI SRE knowledge of your systems</h3><p>When an experienced SRE joins a new team, their skills come with them, but their knowledge of the system doesn't. They spend weeks learning the architecture, the runbooks, and the history of past incidents, along with which services are always noisy. What makes an SRE fast is the combination: years of experience plus deep knowledge of the system they run.</p><p>The same is true for AI agents. LLMs bring some of the experience but none of the knowledge of your system. Without context, an agent spends time and tokens working out how things fit together before it gets to the actual problem.</p><p>The <a href="https://www.elastic.co/search-labs/blog/context-engineering-for-agents">Context Engine</a> gives Elastic nightshift that knowledge before an incident starts. It's the layer that ties the whole system together.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9f88910656b47929/6ac7390182545ed48b57ca0a/image10.png" alt="Context Engine layers: raw telemetry, knowledge indicators, entities and relations, enterprise knowledge, learnings" /><p>The Context Engine builds its knowledge in three ways:</p><ul><li><p><strong>Ahead of time.</strong> As soon as it has access to your data and systems, it extracts knowledge from your telemetry; that is, the entities in your environment, such as services and hosts, and the relationships between them. It can also read your wikis, runbooks, and code from systems like GitHub.</p></li><li><p><strong>From its own work.</strong> Each investigation and detection adds to a long-term memory of what happened and what fixed it.</p></li><li><p><strong>From your interactions and input. </strong>Give it feedback about detections, or feed it a postmortem doc. You can also manually upload important information.</p></li></ul><p>Remembering facts like these is different from the investigation learning described earlier. The Context Engine remembers what's true about your environment, while the Investigation Engine gets better at how it investigates.</p><p>Elastic stores telemetry efficiently enough that you can keep it unsampled, and the agent can search everything within your retention period. Because that telemetry and the knowledge built from it live on Elastic alongside the agent, it can go back to the raw data behind its conclusions.</p><h3>Investigation Engine: AI root cause analysis from any alert</h3><p>Alerts pile up, and each one that matters still lands on an engineer who has to piece together logs, metrics, traces, and recent changes by hand before they can even consider a fix.</p><p>Your new AI SRE coworker takes on that work, using the Investigation Engine. An investigation can start automatically from any alert, whether it’s fired in Elastic or in another tool that you've connected. From there, it does what a good on-call engineer would:</p><ol><li><p>Checks what changed.</p></li><li><p>Forms hypotheses about the cause.</p></li><li><p>Tests them in parallel against your logs, metrics, traces, and profiles.</p></li><li><p>Follows the leads that the evidence supports and drops the ones it doesn't.</p></li><li><p>Separates symptoms from causes.</p></li><li><p>Converges on the most likely root cause.</p></li><li><p>Proposes a fix.</p></li></ol><p>The core logic in our Investigation Engine comes from <a href="https://ir.elastic.co/News--Events/news/news-details/2026/Elastic-Completes-Acquisition-of-Deductive-AI/default.aspx">Deductive AI, which Elastic acquired</a>, and builds on Deductive's three years of work on automated incident investigation.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd2121e2cc3de321e/6ac73933c6656641987d5b9c/image1.png" alt="AI root cause analysis: six hypotheses tested and rejected before the Elastic nightshift AI SRE reaches a conclusion" /><p>The nightshift agent plans a step and runs the tools to carry it out. It observes what comes back and checks its conclusions against the evidence before moving on. Across investigations, it learns which paths lead to useful evidence and successful outcomes, so it gets better at investigating with every incident. </p><p>We talked earlier about building trust when using agentic systems, and this is one example of earning it from you. You can expand any step of an investigation to see the reasoning behind it and the queries that the agent ran. You can also see the raw data that it ran them against.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4bdbb1f7e6a4e86c/6ac73956d1edc68b41573932/image4.png" alt="Expanded hypothesis in an AI SRE investigation showing confidence, the tools it ran and the charts behind the result" /><p>When an investigation finishes, you get:</p><ul><li><p><strong>A conclusion.</strong> The root cause, in plain language, and the immediate fix that we’d recommend.</p></li><li><p><strong>Impact.</strong> Which services are affected and how many problems each one has.</p></li><li><p><strong>Proposed actions.</strong> Next steps, each waiting for your review.</p></li><li><p><strong>Blind spots.</strong> What the agent couldn't verify, so you know where its confidence ends.</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt55636698924d5624/6ac7397e512a6d2599e35d81/image5.png" alt="Completed AI SRE investigation showing what happened, impact across three services and the confirmed root cause" /><h3>Detection Engine: How do you catch incidents with no alert rule?</h3><p>The hardest incidents are often the ones that never triggered an alert to begin with. As a result, the service degrades and users notice. Monitoring stays quiet. No one can write an alert rule for a failure that they haven't imagined, and when something does fire, the related signals are scattered across services, each raising its own alert.</p><p>Your new coworker also takes on that work, using the Detection Engine, which continuously evaluates your telemetry and flags what needs attention. You don’t have to write or maintain any alert rules. When related signals show up across several services, it groups them into a single significant event that shows the signals behind it and the parts of your system that it affects. Significant Events can start an investigation automatically, if you choose, so that it’s already underway by the time someone looks.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7b6e111558a36739/6ac7399ec66566737c7d5ba0/image2.png" alt="Detection Engine grouping error spikes into a Significant Event, with the ES|QL query and sample logs behind it" /><h3>Remediation Engine: Automated remediation from the diagnosis</h3><p>Historically, Observability often stopped at the diagnosis. It could tell you that something was wrong, and eventually why, but the fix was up to you. Runbook automation could only replay the fixes that someone had scripted in advance, and real incidents rarely follow the script.</p><p>Your same new coworker also takes on this work, using the Remediation Engine, which recommends a fix. It can pick one of your existing workflows or connector actions in Elastic Workflows and fill in the parameters for this incident, or it can give you commands or instructions that you can copy into a terminal or a coding agent. Unlike a runbook script, the action and its parameters come from the diagnosis of the incident in front of it. You decide how much it does on its own; you review each action and the parameters it chose before it runs, or let it apply fixes automatically.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt517456f01cbbc68c/6ac739bdd056023fad294ec4/image7.png" alt="Automated remediation: proposed workflow with confidence scores waiting for review before the AI SRE runs it" /><h2>Using the AI SRE from Kibana, your terminal, or Slack</h2><p>Kibana will always be a first-class home for Elastic nightshift, and we'll keep making it a great experience. The Elastic nightshift home page lists your team’s open and recent investigations, sorted by severity, so whoever is on call can see at a glance what needs attention and what's already been handled.</p><p>However, more and more engineers want to use our products from the tools they already live in, like a terminal, a coding agent, or a chat channel, and we're embracing that:</p><ul><li><p><strong>Elastic CLI.</strong> You'll be able to pick up an investigation from your terminal or from a coding agent, like Claude, Cursor, or Codex, and carry on from there. Where Model Context Protocol (MCP) is the easier way to connect an agent, you can use that instead.</p></li><li><p><strong>Slack.</strong> We're building a Slack integration that makes Elastic nightshift feel like a coworker in your channels. It will post finished investigations (and answer questions about them) where your team already talks. More importantly, it will learn from any interaction in Slack. Nudge it along with additional context, or post a root cause analysis or postmortem document in the Slack thread to automatically trigger a learning loop that will use all of this information to get better for the next incident. </p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte113463856bfd07c/6ac739df4ad07f826b385d92/image6.png" alt="Elastic nightshift AI SRE posting a correlated investigation and proposed fix in a Slack incident channel" /><h2>Join the waitlist for the AI SRE private preview</h2><p>Elastic's own SRE team already runs Elastic nightshift against our production environment, and what team members learn goes straight back into the product. With the private preview, we're opening it to more teams.</p><ul><li><p><strong>Who can join:</strong> Existing Elastic users, at no cost during the early access period.</p></li><li><p><strong>What's included:</strong> A preview of the four engines described above, with more capabilities arriving every week.</p></li><li><p><strong>What we ask:</strong> Honest feedback. We'll work directly with preview teams, and what you tell us will shape what we build next.</p></li><li><p><strong>What's next:</strong> Technical preview in Elastic Cloud Serverless later this year.</p></li></ul><p><a href="https://www.elastic.co/observability/ai-sre/signup">Sign up for the private preview</a></p>]]></content:encoded>
    <link>https://www.elastic.co/observability-labs/blog/ai-sre-elastic-nightshift</link>
    <guid isPermaLink="false">ai-sre-elastic-nightshift</guid>
    <category><![CDATA[Agentic Observability]]></category>
    <category><![CDATA[Incident Management]]></category>
    <category><![CDATA[AI]]></category>
    <dc:creator><![CDATA[Luca Wintergerst]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd195024f55f200a2/6ac737f43022640353bb5887/image9.png" length="0" type="image/png"/>
    <pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[From raw logs to system knowledge: the AI context layer observability is missing]]></title>
    <description><![CDATA[A self-updating knowledge base built from your logs: services, dependencies, and failure modes, so your AI agents always know what they are looking at.]]></description>
    <content:encoded><![CDATA[<p>Your monitoring system sees everything, but understands almost nothing.</p>
<p>Before you can rely on most tools to trigger a meaningful alert, you have to do the heavy lifting of telling them exactly what to watch. You have to write the rules, specify what a "normal" baseline looks like, and manually define your service catalog. We're working to change that dynamic at Elastic, and the first major building block is now in place: a system designed to simply read your logs and figure out what's inside them on its own.</p>
<p>Consider what happens when an alert fires today. Your on-call engineer opens an investigation, and the first few minutes are inevitably burned reconstructing basic facts. They have to figure out which services are involved, how those services connect to one another, what error patterns are typical, and which queries they actually need to run to dig deeper. An AI agent faces this exact same cold start problem. Without prior knowledge of your system's architecture, an agent has to read through hundreds of log lines just to establish baseline context that really should already be available.</p>
<p>This blank slate is the default state of most observability setups. You only know what you've explicitly configured. When new services spin up and start writing logs, they sit there without rules until someone takes the time to write them. When architectural dependencies shift, your topology map quietly goes stale unless you've done an exceptional job instrumenting all your services. If an error pattern fires every day but nobody wrote a specific rule to catch it, it remains invisible.</p>
<p>Knowledge Indicators (KI) are our way of closing this gap. When you run extraction against a log stream, Elastic analyzes the raw data and returns structured facts about your environment. It identifies which services are running, the underlying infrastructure they rely on, how they depend on each other, and the log schemas they're using. It even generates a set of ES|QL queries for conditions that might be worth alerting on. Rather than a static configuration, this knowledge accumulates over time, automatically expires when a service disappears, and feeds directly into downstream capabilities like Rules, topology maps, AI agent investigations, and dashboards.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltff3305086ef2e39b/6a7f07c3bdcff054f3c42c2f/topology-graph@2x.png" alt="Knowledge Indicators graph: service nodes (claim-intake, fraud-check, policy-lookup, payment-processor, kafka, notification-dispatch, kubernetes) connected by labeled dependency edges (connection refused, ECONNREFUSED, gRPC UNAVAILABLE, pool exhausted, pod sync)" />
<em>Topology graph generated from dependency KIs: service nodes, dependency edges, and detected error conditions.</em></p>
<h2 id="theextractionpipeline">The Extraction Pipeline</h2>
<p>When designing this system, our primary goal was to eliminate the need for prior context. There should be no mandatory schemas, no service catalogs tied to specific properties, and no predefined static assets that would need to be maintained. We asked ourselves a simple question: if you handed a sample of raw logs to an engineer who had never seen the system before, what could they deduce just by looking?</p>
<p>That thought experiment became our core approach. The system samples a small batch of logs from a stream, processes them through a combination of LLM analysis and deterministic code generators, and accumulates its findings across multiple rounds, entirely configuration-free.</p>
<p>Imagine hiring a room full of junior SREs with one specific job: read these log lines and report their observations, not to fix anything or trigger alarms, just to notice things. "This looks like an nginx server," or "This database is PostgreSQL," or "Service A is calling Service B over HTTP." That's essentially what our extraction job is doing continuously across your streams.</p>
<p>To see how this works in practice, take a look at this single line from an nginx access log:</p>
<pre><code>192.168.1.45 - - [31/Mar/2026:14:23:01 +0000] "POST /api/v2/claims HTTP/1.1" 200 1247 "-" "claim-intake/1.4.2"
</code></pre>
<p>From just this string, the pipeline extracts three distinct facts:</p>
<ul>
<li><strong>Entity</strong>: <code>claim-intake</code> (identifiable as a service from the User-Agent)</li>
<li><strong>Version</strong>: <code>1.4.2</code> (extracted from the User-Agent string)</li>
<li><strong>Technology</strong>: nginx (the web server fielding the request)</li>
<li><strong>Schema</strong>: Combined Log Format</li>
</ul>
<p>Similarly, consider this Java service log:</p>
<pre><code>2026-03-31T14:23:03.412Z INFO fraud-check --- [nio-8080-exec-3] c.e.FraudCheckService : Calling upstream POST http://policy-lookup:8081/v1/policy latency=142ms status=200
</code></pre>
<p>Here, the extraction identifies:</p>
<ul>
<li><strong>Entity</strong>: <code>fraud-check</code> (a Spring Boot service)</li>
<li><strong>Dependency</strong>: <code>fraud-check</code> → <code>policy-lookup</code> (via an outbound HTTP call)</li>
<li><strong>Technology</strong>: Java, Spring Boot</li>
</ul>
<p>Pull twenty lines like these from across your stream, and you quickly build a working, accurate picture of your system architecture.</p>
<p>To ensure this process never blocks ingestion, extraction runs entirely as a background task. You can trigger it on demand from the stream detail view or the Significant Events Discovery UI, but the goal is to have it running by default without requiring attention.</p>
<p>The pipeline itself runs multiple iterations, each time fetching a small sample of documents. We use a mix of random and already-excluded documents to ensure we discover the full scope of the system. KIs found in one iteration are fed back as exclusions into the next, so each round focuses on what the previous one missed—ensuring quieter, less-represented services aren't crowded out by noisier ones.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdfa9e57d3cb5b3bc/6a7f07c6b43770fb324d6a9d/ki-extraction-pipeline@2x.png" alt="KI extraction pipeline: raw logs → three-pool biased sampling (entity-filtered / diverse / random) → LLM finalize_features and 4 computed generators in parallel → merge and dedup → 84 KIs stored" />
<em>Extraction pipeline: biased document sampling feeds a parallel LLM pass and four deterministic generators. Results are merged and deduplicated before storage.</em></p>
<p>Once sampled, the documents are sent to an LLM. We use a system prompt that instructs the model to identify a few specific types of features, which we plan to extend over time:</p>
<p>| Type | What it captures |
|------|------------------|
| Entity | Distinct system components: services, applications, jobs |
| Infrastructure | Environment context: Kubernetes, cloud provider, OS |
| Technology | Languages, frameworks, libraries, databases |
| Dependency | Relationships between components |
| Schema | Log format conventions: ECS, OTel, custom |</p>
<p>The LLM returns its findings, delivering newly identified traits alongside any intentionally ignored ones (like user-excluded false positives). To be accepted, every feature must include stable identifying properties and cite direct evidence from the sampled logs. The LLM also assigns a confidence score from 0–100 for each KI, so any downstream use of that KI knows how much to trust it.</p>
<p>In parallel, a set of deterministic code-based generators independently analyze the data to produce statistical summaries, log samples, pattern clusters, and error-specific features. Because these are computed rather than inferred, they always receive a confidence score of 100.</p>
<p>Finally, the LLM results and computed features are merged and deduplicated. Known KIs reuse their existing UUIDs, new discoveries get fresh ones, and any user-excluded features are quietly dropped server-side. Surviving KIs are saved with an active status and an expiration date set for seven days out.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1aefe63628f536ce/6a7f07c9fc63ab916364ca61/kis.png" alt="Knowledge Indicators tab showing 84 KIs across streams, with type, confidence (1–5 stars), and stream columns visible" />
<em>Knowledge Indicators tab showing 84 KIs across streams, with type, confidence (1–5 stars), and stream columns.</em></p>
<h2 id="whataknowledgeindicatorcontains">What a Knowledge Indicator Contains</h2>
<p>Knowledge Indicators fall into two categories: Feature KIs and Query KIs.</p>
<p>Feature KIs are descriptive. They explain the contents of the stream: what services are running, the infrastructure housing them, their dependencies, and the active tech stack.</p>
<p>Query KIs are actionable. They are ready-to-run ES|QL queries targeting notable conditions like connection exhaustion, out-of-memory errors, or fatal exceptions. Each comes with a severity score from 0 to 100, and when promoted to Rules, they fire Events.</p>
<p>Feature KIs carry a full data model:</p>
<ul>
<li><strong><code>type</code> / <code>subtype</code></strong>: the category of the fact (Entity, Infrastructure, Technology, Dependency, Schema)</li>
<li><strong><code>title</code> / <code>description</code></strong>: a human-readable summary</li>
<li><strong><code>properties</code></strong>: stable key-value pairs used to deduplicate findings across multiple runs</li>
<li><strong><code>confidence</code></strong>: 0–100. LLM-identified KIs score based on evidence quality. Deterministic KIs always score 100.</li>
<li><strong><code>evidence</code></strong>: 2–5 supporting log excerpts that justify the KI's existence</li>
<li><strong><code>filter</code></strong>: an optional StreamLang condition scoping the KI to specific documents</li>
</ul>
<p>A dependency KI looks like this:</p>
<pre><code>{
  "type": "dependency",
  "subtype": "service_dependency",
  "title": "api_gateway → inference_service",
  "description": "Service-to-service HTTP dependency from api_gateway to inference_service, observed in request logs",
  "properties": {
    "source": "api_gateway",
    "target": "inference_service",
    "protocol": "http"
  },
  "confidence": 85,
  "evidence": [
    "service.name=api_gateway http.url=/v1/inference peer.service=inference_service",
    "upstream=inference_service:8080 request=POST /v1/inference 200"
  ],
  "filter": { "field": "service.name", "eq": "api_gateway" },
  "status": "active",
  "expires_at": "2026-04-09T00:00:00Z"
}
</code></pre>
<p>Query KIs take a simpler shape, focusing solely on the title, severity score, and the executable query:</p>
<pre><code>{
  "kind": "query",
  "title": "PostgreSQL connection slot exhaustion",
  "description": "Fires when Postgres runs out of available connection slots",
  "severity_score": 90,
  "esql": {
    "query": "FROM logs-* | WHERE service.name == \"postgres\" AND message : \"remaining connection slots\""
  }
}
</code></pre>
<p>The <code>properties</code> field is what keeps Feature KIs stable across multiple pipeline runs. The dependency KI for <code>api_gateway → inference_service</code> records the source, target, and protocol as fixed pairs. The next time extraction runs, Elastic recognizes this existing relationship and updates the KI's <code>last_seen</code> timestamp rather than creating a duplicate.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcf99c981e914aefc/6a7f07cc6693f89c62663d5b/ki-detail.png" alt="KI detail panel for api_gateway → inference_service showing type, subtype, properties, confidence, evidence, and expiry date" />
<em>KI detail panel showing type, subtype, properties, confidence, evidence, and expiry date for a service dependency.</em></p>
<h2 id="thefoundationforintelligentobservability">The Foundation for Intelligent Observability</h2>
<p>So what can we do with all of this? These KIs serve as the contextual foundation for Elastic's more advanced capabilities. From just these extracted KIs, we can automatically generate active Rules to surface interesting signals, without a human engineer writing a single line of configuration. More on this particular capability in the next post in this series.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2ab47e653ddc6683/6a7f07d01967eada7c330527/queries.png" alt="85 auto-generated Rules from KIs with impact ratings (Critical/High) and event occurrence sparklines" />
<em>85 auto-generated Rules from 84 KIs, with impact ratings and event occurrence sparklines.</em></p>
<p>As a user or an agent, the dependency KIs automatically construct an infrastructure graph—inferred entirely from log data, not from distributed tracing or any manual configuration. During an incident, this graph is invaluable for assessing blast radius. If a specific database goes down, the topology map immediately shows you exactly which upstream services are about to fail, without maintaining a manual service catalog.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb87332c1b105dd39/6a7f07d373d9bd5d0b29d946/topology-dependencies@2x.png" alt="Topology graph showing service dependencies: claim-intake connects to claim-intake-db (PostgreSQL), fraud-check, policy-lookup; fraud-check connects to fraud-check-db (MongoDB); policy-lookup connects to policy-lookup-db (PostgreSQL); payment-processor and kafka grouped separately; notification-dispatch and kubernetes at bottom" />
<em>Service dependency graph extracted from KIs, showing services, databases, and infrastructure components.</em></p>
<p>This context changes how an AI agent handles an incident. Instead of starting from scratch, the agent initiates its investigation using your system's actual topology and known failure modes. Based on the KIs, it identifies the relevant streams, runs the applicable queries, and formulates a specific hypothesis. In our example, it already knows that <code>api_gateway</code> relies on <code>inference_service</code>, and it knows that connection slot exhaustion is a high-severity failure mode for your Postgres instance.</p>
<p>This extracted knowledge doesn't have to be perfect to be useful. Because LLMs are inherently non-deterministic, a KI might occasionally be slightly off, but it still gives the agent a significant head start. The agent can cross-reference the KI against live logs and self-correct on the fly. The real benefit is simply not having to reconstruct basic facts during a critical outage. KIs also drive AI-generated dashboard suggestions and inform Grok pattern generation whenever you introduce new streams.</p>
<h2 id="selfcleaningandscalable">Self-Cleaning and Scalable</h2>
<p>Maintaining this knowledge base is entirely hands-off. KIs auto-expire after 7 days if they aren't observed in subsequent extraction runs. If you decommission a service, its associated KIs simply fade away without any manual cleanup. If the service comes back online later, the KIs are re-extracted. Users can also mark individual feature KIs as false positives, and the system carries those exclusions forward into future runs to prevent re-identification.</p>
<p>Because we scoped KI extraction as a specific classification task, looking at around 20 log samples to identify services, infrastructure, and dependencies, it doesn't require a large frontier model to run. A fast, cost-effective model handles this without multi-step reasoning.</p>
<h2 id="youshouldnthavetotellyourtoolswhattowatch">You Shouldn't Have to Tell Your Tools What to Watch</h2>
<p>The fundamental promise of observability is to help you understand your systems. For far too long, the burden of teaching the tool how those systems actually work has fallen on the engineers operating them.</p>
<p>The next post in this series looks at what agents do with that context: why every agent that investigates your system without KIs re-learns the same things from scratch on every incident, and what changes when it doesn't have to.</p>
<p><strong>NOTE:</strong> These capabilities are available behind a feature flag in Serverless Observability projects. Turn on <code>observability:streamsEnableSignificantEvents</code> by searching for it in the Kibana advanced settings page. </p>
<h2 id="frequentlyaskedquestions">Frequently asked questions</h2>
<p><strong>What are Knowledge Indicators in Elasticsearch Streams?</strong>
Knowledge Indicators (KIs) are structured facts extracted from raw log streams: service names, infrastructure components, service-to-service dependencies, and tech stack details. Elastic extracts them automatically by sampling log lines, without requiring schemas, service catalogs, or manual configuration.</p>
<p><strong>How does Elastic build a service topology map from logs alone?</strong>
The extraction pipeline samples log lines and identifies dependency relationships, such as an outbound HTTP call from one service to another. These dependency KIs are used to construct a topology graph that shows which services depend on which, entirely inferred from log data, without distributed tracing or any manual input.</p>
<p><strong>Why does an AI agent need Knowledge Indicators before investigating an incident?</strong>
Without KIs, an AI agent starts every investigation from scratch: it has to read hundreds of log lines just to establish which services exist and how they relate. KIs give the agent a pre-built map of your system, including services, known failure modes, and relevant queries, so it can begin reasoning about the actual incident immediately.</p>
<p><strong>Do I need to configure anything for Knowledge Indicator extraction to work?</strong>
No. The pipeline requires no schema definitions, no service catalog, and no predefined rules. It samples a small set of log lines from a stream, analyzes them through a combination of LLM inference and deterministic generators, and accumulates findings automatically.</p>
<p><strong>How accurate are LLM-extracted KIs compared to computed ones?</strong>
Computed (deterministic) KIs always receive a confidence score of 100 because they are derived from statistical analysis rather than inference. LLM-extracted KIs receive scores from 0 to 100 based on the quality of evidence found in the sampled logs. Rules, agent investigations, and topology maps can all use this score to weight their decisions.</p>
<p><strong>What happens when a service is decommissioned?</strong>
KIs carry a 7-day expiration. If a service stops appearing in subsequent extraction runs, its KIs expire and are removed automatically. No manual cleanup required. If the service comes back, the KIs are re-extracted on the next run.</p>
<p><strong>How does this compare to service discovery via distributed tracing?</strong>
Distributed tracing requires instrumented services and a trace collector. Knowledge Indicator extraction requires nothing beyond existing log streams: no SDK, no agent, no schema. For environments with partial or no tracing coverage, KI extraction provides topology and dependency information that tracing would otherwise miss.</p>]]></content:encoded>
    <link>https://www.elastic.co/observability-labs/blog/elastic-knowledge-indicators-log-extraction</link>
    <guid isPermaLink="false">elastic-knowledge-indicators-log-extraction</guid>
    <category><![CDATA[Machine Learning]]></category>
    <category><![CDATA[Logs Analytics]]></category>
    <dc:creator><![CDATA[Luca Wintergerst]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4563c156c7dfcd41/6a7f07d66c6eac3466f13f0b/cover.png" length="0" type="image/png"/>
    <pubDate>Tue, 05 May 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Fixing Elastic Streams processing failures without dropping data]]></title>
    <description><![CDATA[When your Streams ingest pipeline breaks, failed documents land in the failure store, not the floor. Here's how to use those exact failures to fix your pipeline without re-ingesting from the source.]]></description>
    <content:encoded><![CDATA[<p>If you've run a Streams pipeline for more than a week, you've probably hit a processing failure. Before Streams, that often meant dropped data or a dead letter queue at the shipper layer: extra infrastructure you had to operate separately. Here's the recovery loop today.</p>
<h2 id="whenprocessingfailsdatalandsinthefailurestore">When processing fails, data lands in the failure store</h2>
<p>When a Streams pipeline fails (a Grok pattern doesn't match, a field type conflicts with the mapping), the documents that caused the failure are written to the <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/failure-store">failure store</a>. The failure store is a set of backing indices attached to your data stream. It scales the same way as any other data stream, so it can absorb everything that fails. It's enabled by default for logs as of Elasticsearch 9.2.</p>
<p>The <strong>Data quality</strong> tab gives you insights into the quality of your stream and into documents in the failure store. When failures are accumulating, you'll see a rising count of failed documents along with the error type and a sample of the messages that triggered it.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4686e6817c90d63e/6a7f09635967e508545dd162/processing-failures-in-failure-store.png" alt="Processing failures accumulating in the failure store" />
<em>The Data quality tab showing a rising failure count, error type, and a sample of the documents that triggered it.</em></p>
<p>A Grok expression mismatch (<code>illegal_argument_exception</code>) is sending documents to the failure store. The raw log line doesn't match the expected pattern. The documents aren't dropped. They're in the failure store, ready to debug against.</p>
<h2 id="processingswitchthesamplesourcetothefailurestore">Processing: Switch the sample source to the failure store</h2>
<p>Start by navigating to the <strong>Processing</strong> tab.</p>
<p>By default, the editor samples from recent live documents. Switch the sample source to <strong>Failure store</strong> instead: it loads the exact documents that failed, the unmodified originals before any Streams processing ran. You're iterating against the actual failures.</p>
<p>Change the sample source dropdown from the default to Failure store.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta92cf204670630eb/6a7f09666693f82c81663def/sample-source-dropdown.png" alt="Sample source dropdown showing Latest samples and Failure store options" />
<em>The sample source dropdown with the Failure store option selected.</em></p>
<p>The editor loads up to 100 documents from the failure store and runs them through the current pipeline. You can see exactly where parsing breaks down.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt528d61760756ab0f/6a7f09694c4bfb9ca2ccd3d9/failure-store-samples-processing.png" alt="Pipeline editor with failure store selected as the sample source" />
<em>The pipeline editor loaded with documents from the failure store instead of recent live samples.</em></p>
<h2 id="fixtheprocessoragainsttheactualfailures">Fix the processor against the actual failures</h2>
<p>With the failure store documents loaded as samples, iterate on the processor. The editor shows you the result against the actual failed documents in real time.</p>
<p>In this example, the pipeline was originally built to parse HTTP access logs:</p>
<pre><code>DELETE /api/v1/auth/logout from 26.72.241.177 - Status: 200 - Response time: 38ms - Request ID: req_24363339 - Location: São Paulo, BR - Device: desktop
HEAD /api/v1/notifications from 20.94.145.254 - Status: 202 - Response time: 60ms - Request ID: req_74513322 - Location: Tokyo, JP - Device: mobile
</code></pre>
<p>The original Grok pattern matched those:</p>
<pre><code>%{WORD:http.method} %{URIPATH:uri.path}
</code></pre>
<p>A second log type started flowing in. Cache hits and external API calls arrived in a different format:</p>
<pre><code>cache_hit: Cache hit for key: config
external_api_call: External API call completed - latency: 1695ms - Duration: 598ms
</code></pre>
<p>The original pattern doesn't match these at all. Every one goes straight to the failure store. With the failure store loaded as the sample source, the problem is immediately obvious: the editor shows the parse failing on lines that start with a word followed by a colon, not an HTTP method followed by a path.</p>
<p>The fix is a second pattern to handle the new format:</p>
<pre><code>%{WORD:event.type}: %{GREEDYDATA:message}
</code></pre>
<p>Add it to the processor, and the editor immediately shows both log types parsing correctly against the failure store samples.</p>
<p>When the sample view shows all fields extracting correctly and the parse rate hits 100%, the fix is ready.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1ffb5b202ba966a8/6a7f096d63e959962d73dc6a/successful-parsing.png" alt="Pipeline editor showing successful parsing against failure store samples" />
<em>Both log types parsing correctly after adding the second Grok pattern. Parse rate at 100%.</em></p>
<p>No guessing — the editor confirms the fix before you save.</p>
<h2 id="watchthefailurecountdrop">Watch the failure count drop</h2>
<p>Save the updated pipeline. New documents are now processed with the corrected pipeline. Switch back to the Data quality tab and watch the failure count.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6fd207484c0a8170/6a7f096f73d9bdcc0129d9c9/resolved.png" alt="Failure store count dropping after pipeline fix" />
<em>The failure count dropping as new documents are processed by the corrected pipeline.</em></p>
<p>The count drops as the fixed pipeline handles new incoming data correctly. The remaining documents in the failure store are the pre-fix failures. They'll clear out as retention ages them off.</p>
<p>The fix applies to new documents only. Documents already in the failure store aren't automatically reprocessed; each was processed by the pipeline version active when it arrived. If you need them in your main stream, that's a separate step.</p>
<h2 id="therecoveryloop">The recovery loop</h2>
<p>Open Data quality, switch to the failure store, fix the processor, save. The whole thing takes a few minutes at most.</p>
<p>No re-ingestion from source. No shipper-level dead letter queue to operate. If you haven't checked the Data quality tab for your streams recently, it's worth a look. There might be failures sitting there that a one-line fix would clear.</p>
<p>For a deeper look at what the Data quality tab shows and how to configure the failure store, see <a href="https://www.elastic.co/observability-labs/blog/data-quality-and-failure-store-in-streams">Elastic Observability: Streams Data Quality and Failure Store Insights</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/observability-labs/blog/elastic-streams-failure-store-processing</link>
    <guid isPermaLink="false">elastic-streams-failure-store-processing</guid>
    <category><![CDATA[Logs Analytics]]></category>
    <dc:creator><![CDATA[Luca Wintergerst]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt4686e6817c90d63e/6a7f09635967e508545dd162/processing-failures-in-failure-store.png" length="0" type="image/png"/>
    <pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch over the years — how LogsDB cuts index size by up to 75% at no throughput cost]]></title>
    <description><![CDATA[By default, Elasticsearch is optimized for retrieval, not storage. LogsDB changes that. Here's the layered architecture behind a 77% index size reduction.]]></description>
    <content:encoded><![CDATA[<p>Elasticsearch was built as a search engine. That heritage has a cost for log storage: every event fans out to multiple on-disk structures, each optimized for retrieval rather than compression. LogsDB changes both. On our nightly benchmark, Enterprise mode produces a 37.5 GB index from the same data that takes 161.9 GB without LogsDB — a 77% reduction from a single setting.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta3a4d65c4de793d4/6a7f09e73cab1cd73c0e4734/storage-breakdown-v3-bold@2x.png" alt="Standard vs LogsDB storage breakdown" /></p>
<h2 id="thewriteoverhead">The write overhead</h2>
<p>Lucene, the library underneath, keeps multiple structures for every indexed document:</p>
<ul>
<li>The <strong>inverted index</strong> maps terms to documents. This is what makes text search fast.</li>
<li><strong><code>_source</code></strong> stores the original JSON blob, returned when you fetch a document.</li>
<li><strong>Doc values</strong> store field values in columns for sorting and aggregation.</li>
<li><strong>Points / BKD trees</strong> index numeric and date fields for range queries.</li>
</ul>
<p>The inverted index earns its keep: it's what lets you search a billion log lines by keyword in milliseconds, and there's no cheaper way to build that capability. <code>_source</code> exists to give you back exactly what you indexed: search results and <code>GET</code> requests return this blob directly. The problem is that it stores the full event even though the same field values are already available through doc values and the other structures.</p>
<p>Take a log event with fields like <code>host.name</code>, <code>@timestamp</code>, <code>http.response.status_code</code>, and <code>duration_ms</code>. The entire event is serialized as JSON in <code>_source</code>. The same field values are also written into doc values columns, indexed into the inverted index, and stored in BKD trees for range queries. Same data, multiple structures, each with its own on-disk footprint.</p>
<p>For a search engine where you need fast retrieval across all dimensions, that overhead is a reasonable tradeoff. For logs, where you rarely need the raw JSON and almost never do relevance-ranked search, much of it is pure waste.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltbf0f50c8d9b1c3c5/6a7f09e963e95944bd73dca6/dual-storage-bold@2x.png" alt="One incoming log event fans out to four on-disk structures" />
<em>One write, four on-disk structures: <code>_source</code> (the raw JSON blob), the inverted index, doc values columns, and BKD / points trees for numeric range queries. The same field values end up in multiple places.</em></p>
<h2 id="whycolumnarstoragemattersforcompression">Why columnar storage matters for compression</h2>
<p>Doc values are the key to everything LogsDB does. Unlike <code>_source</code>, which stores entire documents as blobs, doc values store each field as a separate column across all documents in a Lucene segment.</p>
<p>Picture a segment with a million log events. The <code>_source</code> representation is a million JSON blobs, one per event, each containing all fields jumbled together. The doc values representation is a set of columns: one column of a million timestamps, one column of a million host names, one column of a million status codes, and so on.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta2548bca8f377f56/6a7f09eceab5bec47820a589/doc-values-columns-bold@2x.png" alt="Row-oriented vs column-oriented storage" />
<em>Row-oriented <code>_source</code> keeps all fields for each document in one blob — doc0 through doc5 each carry <code>host.name</code>, <code>@timestamp</code>, <code>status</code>, <code>duration_ms</code>, and more jumbled together. Column-oriented doc values restructure the same data so all <code>host.name</code> values sit in one column, all timestamps in another, all status codes in another. Compression codecs can then run on each contiguous column independently.</em></p>
<p>That columnar layout is what makes per-column compression possible. When all values of <code>http.response.status_code</code> sit in a contiguous column, Lucene can apply codecs that exploit patterns in the sequence.</p>
<p>Delta encoding stores differences between adjacent values instead of full values. GCD encoding finds a common factor and divides everything down. Run-length encoding collapses repeats. Lucene picks the codec per segment and re-evaluates when segments merge.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltcab1e02d2192a98b/6a7f09efb6b7345f6ae48cc6/numeric-codec-pipeline-bold@2x.png" alt="Numeric codec pipeline: RAW → DELTA → GCD → BIT-PACK" />
<em>Four sorted <code>@timestamps</code> from the same host, compressed in four stages. RAW: four 32-bit integers, 128 bits total. DELTA: store differences instead of full values — base stays, deltas +100, +200, +300 take 59 bits. GCD: divide out the common factor of 100, leaving 1, 2, 3 at 39 bits. BIT-PACK: pack those three small integers into contiguous bit storage, 9 bits freed.</em></p>
<p>But here's the catch: these codecs only work well when adjacent documents have correlated values. Consider the <code>@timestamp</code> column.</p>
<p>If logs arrive from dozens of hosts interleaved randomly, the timestamps in the column jump around. The delta between adjacent values might be +3 seconds, then -47 seconds, then +120 seconds. Delta encoding can't do much with that.</p>
<p>Now consider what happens if you sort by <code>host.name</code> and <code>@timestamp</code> before writing to the segment. All logs from host-A land in a contiguous run, followed by all logs from host-B, and so on. Within each host's run, the timestamps are monotonically increasing and the deltas are predictable.</p>
<p>Four timestamps from the same host might look like 1706745600, +100s, +200s, +300s. Delta encoding shrinks those to a base value plus three small integers.</p>
<p>GCD encoding finds that 100, 200, 300 are all divisible by 100 and stores 1, 2, 3 instead. Bit-packing then fits those three values into a handful of bits. The same pattern applies to fields like <code>host.name</code>, <code>service.name</code>, or <code>http.response.status_code</code>: within a sorted run, long stretches of identical values collapse to near nothing under run-length encoding.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1c5e01c48262da24/6a7f09f1de231589e4fd7afd/index-sorting-bold@2x.png" alt="Index sorting: arrival order → sorted by host.name → after RLE" />
<em>Five hosts — api-01, api-02, db-01, web-01, web-02 — scattered randomly in arrival order (left). Sorting by <code>host.name</code> groups them into five contiguous blocks of eight (center). Run-length encoding collapses each block to a single (value, count) pair — 5 pairs stored instead of 40, the remaining slots freed (right).</em></p>
<p>Elasticsearch never sorted by default. Documents landed in arrival order, compressed with DEFLATE. We left a lot on the table.</p>
<h2 id="howwegothere20122026">How we got here: 2012–2026</h2>
<p>Not all of the individual techniques in LogsDB were designed for logs. They were built over twelve years to solve different problems, and LogsDB is what happens when you stack them.</p>
<p><strong>The foundation (2012–2017).</strong> Lucene 4.0 introduced doc values in 2012. By Elasticsearch 5.0 in 2016, they were on by default for all keyword and numeric fields. Lucene 7.0 added sparse doc values, so fields that only appear in some documents don't waste space on every document in the segment. That fixed a significant force-merge bloat problem (up to 10× on sparse fields) and set up the storage model everything else depends on.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte4b2c68318492f3d/6a7f09f4227b1c2e2b5984aa/sparse-doc-values-bold@2x.png" alt="Dense vs sparse doc values encoding" />
<em>Dense encoding reserves an 8-byte slot per document regardless of presence. Sparse encoding stores only documents that have a value at 12 bytes each (value + doc ID). For <code>error_code</code> with 2 of 16 docs populated (12% fill), sparse is 81% smaller: 24 B vs 128 B. For <code>request_path</code> at 88% fill, sparse is larger: 168 B vs 128 B. Lucene picks per field; sparse wins below ~67% fill.</em></p>
<p><strong>Incremental wins (2020–2021).</strong> Two smaller changes targeted observability workloads. Dictionary-based stored fields compression deduplicated repetitive string metadata for about a 10% win.</p>
<p>The <code>match_only_text</code> field type dropped term frequencies and positions from the inverted index. Term frequencies are what BM25 uses to score documents by relevance — how often a term appears in a document relative to the rest of the corpus. For log search that signal is meaningless: you don't care whether "timeout" appeared twice or seven times in a log line, you just want to find it. Positions are similar: they're stored so Elasticsearch can do exact phrase matching, but the position data is expensive and phrase queries on logs are rare enough that the tradeoff is worth it. When you do run a phrase query on a <code>match_only_text</code> field, it still works — it just falls back to a slower path that rescores candidates rather than using stored positions directly.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc4dd1dc99cf753b0/6a7f09f76693f85f60663e1f/match-only-text-bold@2x.png" alt="text vs match_only_text inverted index storage" />
<em><code>text</code> stores each term with its frequency and every position it appears at. <code>match_only_text</code> keeps only the doc IDs — enough to find the document, nothing more. The <code>timeout</code> term appears twice in this message (positions 1 and 4), which is exactly the kind of data that gets dropped.</em></p>
<p>Dropping frequencies and positions cuts the inverted index for a text field by roughly 40%. The overall index impact in 2021 was only ~10%, which sounds like a poor return on a 40% field-level reduction. The reason is where storage was going at the time: <code>_source</code> was stored in full for every document as a raw JSON blob, doc values were uncompressed and unsorted, and nothing was using ZSTD. The <code>message</code> field's inverted index was a small slice of a much larger, poorly-compressed whole. As the next five years of work addressed those other structures, the same 40% field-level savings became a meaningful fraction of a much smaller total.</p>
<p>Neither change was decisive on its own, but they established that log-specific storage optimization was worth pursuing.</p>
<p><strong>The TSDB turning point (April 2023).</strong> This is where the story really starts. We shipped synthetic <code>_source</code> and index sorting for time series metrics in Elasticsearch 8.7.</p>
<p>Synthetic source changes the write-and-read contract. At write time, we skip storing the raw JSON blob entirely. At read time, when a query needs to return the original document, we reconstruct it by reading each field's value out of doc values and stored fields and assembling them back into JSON. The result is functionally equivalent to the original <code>_source</code> (with minor differences like field ordering), but we never stored the blob.</p>
<p>Index sorting groups documents by dimension fields and timestamp before writing to disk. Together, synthetic source and index sorting cut metrics storage by up to 70%.</p>
<p>That result told us something important: the same architecture could work for logs.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta1aaae98650a025b/6a7f09fa3cab1c48250e4740/synthetic-source-bold@2x.png" alt="Standard _source vs synthetic _source" />
<em>Without LogsDB, Elasticsearch writes every log event twice: once as a raw <code>_source</code> blob on disk, once into doc values columns. LogsDB skips the blob entirely. At read time, a <code>GET &lt;index&gt;/_doc/1</code> request gathers field values from doc values and assembles the document on the fly.</em></p>
<p><strong>The TSDB codec (2024).</strong> In 8.13 and 8.14, we built a custom doc values codec with run-length encoding optimized for sorted consecutive values, PFOR-delta encoding, and cyclic ordinal encoding for multi-valued dimensions. The numbers were striking: <code>kubernetes.pod.name</code> doc values dropped from 110 MB to 7.25 MB in one benchmark. We extended coverage to all numeric and keyword types including <code>ip</code>, <code>scaled_float</code>, and <code>unsigned_long</code>.</p>
<p><strong>LogsDB Tech Preview (August 2024).</strong> In <a href="https://github.com/elastic/elasticsearch/pull/108896">8.15</a>, we combined everything into <code>index.mode: logsdb</code>: host-first sorting, synthetic <code>_source</code>, ZSTD compression, and the TSDB numeric codecs. One decision mattered more than expected: sort order. Sorting by <code>host.name</code> first, then <code>@timestamp</code>, delivers up to ~40% storage reduction. Sorting by timestamp first gives ≤10%. The host-first ordering co-locates documents that share field values, which is exactly what the numeric codecs need.</p>
<p><strong>ZSTD and GA (November–December 2024).</strong> In <a href="https://github.com/elastic/elasticsearch/pull/112665">8.16</a>, we switched <code>best_compression</code> from DEFLATE to ZSTD permanently (level 3, blocks up to 2,048 documents or 240 kB, native bindings via Panama FFI on JDK 21+). ZSTD gave us ~12% smaller stored fields and ~14% higher indexing throughput at the same time, which almost never happens. LogsDB went GA in 8.17.</p>
<p>At GA, we claimed up to 65% storage reduction.</p>
<p><strong>Routing and recovery (April 2025).</strong> In 8.18, <a href="https://github.com/elastic/elasticsearch/pull/116687"><code>route_on_sort_fields</code></a> started routing documents to shards by sort field values instead of <code>_id</code>. Without this optimization, Elasticsearch hashes the <code>_id</code> to pick a shard, so logs from the same host scatter across all shards. With routing on sort fields, logs with similar <code>host.name</code> values land on the same shard. This co-locates similar documents at the shard level, not just within segments, adding ~20% storage reduction at a 1–4% ingest penalty. Routing on sort fields requires auto-generated <code>_id</code>.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt91e7d6426bc3bcf1/6a7f09fd9090b02c5584e8bf/shard-routing-bold@2x.png" alt="Shard routing: standard, routed, routed + sorted" />
<em>Data stream <code>.ds-logs-nginx-default-00001</code> with six hosts across three shards. STANDARD (hashed by <code>_id</code>): all host colors scattered randomly. ROUTED (<code>route_on_sort_fields</code>): same-host logs land on the same shard, but remain in arrival order within it. ROUTED + SORTED (host-first sort): each shard contains contiguous blocks of a single host — the combination that lets numeric codecs and RLE reach their full potential.</em></p>
<p>We also <a href="https://github.com/elastic/elasticsearch/pull/119110">switched peer recovery to synthetic source reconstruction</a>, eliminating the duplicate <code>_recovery_source</code> blob. In <a href="https://github.com/elastic/elasticsearch/pull/121049">9.0</a>, <code>logs-*-*</code> indices default to LogsDB.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt6a8f273087631cf2/6a7f0a00ea068d7193f09d4b/recovery-source-bold@2x.png" alt="Index size written: _recovery_source eliminated" />
<em>Nightly synthetic source benchmark, December 2024. Index size written drops 39% — from ~279 GB to ~171 GB — the day peer recovery switches from copying the raw <code>_recovery_source</code> blob to reconstructing documents from doc values.</em></p>
<p><strong>Merge and recovery overhaul: 9.1 (July 2025).</strong> We fully eliminated the recovery source. Peer recovery uses batched synthetic reconstruction, cutting write I/O by ~50% and boosting median indexing throughput ~19% over the 8.17 baseline. We replaced up to four separate doc values merge passes with a single pass, cutting background merge CPU by up to 40%. And we swapped <code>_seq_no</code>'s BKD tree for Lucene doc value skippers, halving <code>_seq_no</code> storage.</p>
<p><strong>pattern_text and Failure Store: 9.2–9.3 (October 2025–February 2026).</strong> In <a href="https://github.com/elastic/elasticsearch/pull/124323">9.2</a>, we shipped <code>pattern_text</code> as a Tech Preview: a new field type that decomposes log messages into static templates and dynamic variable parts. A log line like <code>Session opened for user alice from 10.0.1.42 via TLS</code> gets split into the template <code>Session opened for user {} from {} via TLS</code> (stored once, as a template ID) and the variables <code>alice</code>, <code>10.0.1.42</code> (stored per document). For logs with high template repetition, this cuts message field storage by up to 50%. A companion <code>template_id</code> sub-field lets you sort by template, and the LogsDB setting <code>index.logsdb.default_sort_on_message_template</code> enables this automatically. <code>pattern_text</code> <a href="https://github.com/elastic/elasticsearch/pull/135370">went GA in 9.3</a>.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt0621df09262d5bf7/6a7f0a03c2cc097ec224943a/pattern-text-bold@2x.png" alt="TEXT vs PATTERN_TEXT field type" />
<em>TEXT stores each log message as a full string per document — eight copies of near-identical blobs. PATTERN_TEXT decomposes them: the shared template <code>Session opened for user {} from {} via TLS</code> is stored once with ID T0, and only the variable columns (<code>user</code>, <code>ip</code>) are stored per document — alice/10.0.1.42, bob/10.0.1.87, carol/10.0.2.11, and so on.</em></p>
<p><code>pattern_text</code> does come with an indexing CPU cost: decomposing each message into template and variables takes more work at write time than storing a raw string. Whether that tradeoff makes sense depends on your dataset and your priorities.</p>
<p>If your log messages follow highly repetitive patterns (structured application logs, Kubernetes events, access logs), the storage wins are large and the CPU overhead is bounded. If your messages are free-form or low-repetition, the compression gains shrink while the CPU cost stays roughly the same.</p>
<p>For data you keep for months or years, the cumulative storage reduction usually makes it worthwhile. For high-cardinality, rapidly changing messages where storage isn't the constraint, it may not be.</p>
<p>9.3 also brought compression for binary doc values, making <code>wildcard</code> field types significantly more storage-efficient. Internally, wildcard fields store an inverted index of trigrams in a binary doc values column; that column is now compressed with Zstandard instead of being stored raw. In one benchmark, a URL field dropped from 2.92 GB to 1.12 GB, more than 60% compression. If you use <code>wildcard</code> fields heavily, the gain is automatic with no mapping changes needed.</p>
<p>Also in 9.3, skip lists for <code>@timestamp</code> and <code>host.name</code> became available as an opt-in for LogsDB. Skip lists let Elasticsearch jump ahead in a doc values column without reading every entry, which speeds up time-range queries on large segments. Other index modes have skip lists disabled by default; in LogsDB you can enable them selectively for the fields you range-query most.</p>
<p>Also in 9.3, the <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/failure-store">Failure Store</a> <a href="https://github.com/elastic/elasticsearch/pull/131261">became enabled by default</a> for <code>logs-*-*</code> data streams. Failed documents (mapping conflicts, ingest pipeline errors) now land in dedicated <code>::failures</code> indices instead of being rejected, which means LogsDB's strict synthetic source requirements are less likely to cause silent data loss during migration.</p>
<h2 id="performancenotjuststorage">Performance, not just storage</h2>
<p>LogsDB started as a storage optimization, and the early releases came with a throughput cost — sorting, synthetic source reconstruction, and ZSTD all add work at write time. Over two years of releases, we clawed that back. Indexing throughput is now on par with what users had before enabling LogsDB. You get the storage reduction without giving up the ingest rate you were used to.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7d5f55de5dbef47b/6a7f0a06bd2198e809757fa1/performance-over-time-bold@2x.png" alt="LogsDB throughput and storage on disk over time" />
<em>Throughput (teal) has climbed from ~25k to ~35k docs/s since the Tech Preview. Storage on disk (blue) has dropped from ~65 GB to ~36 GB on the same benchmark dataset. Both curves move in the right direction, driven by the same layered releases: ZSTD in 8.16, routing optimization in 8.18, the merge and recovery overhaul in 9.1. Live numbers at <a href="https://elasticsearch-benchmarks.elastic.co/#tracks/logsdb/nightly/default/90d">elasticsearch-benchmarks.elastic.co</a>.</em></p>
<p>The two trends compound each other. Less storage means fewer segments to merge, which frees CPU for indexing. Synthetic source reconstruction is cheaper to compute than it is to store and replicate the raw blob. Each release that shrank the index also reduced background I/O, which fed back into throughput.</p>
<p>The practical result: if you were running standard Elasticsearch for log ingestion two years ago, the throughput you had then is roughly what LogsDB delivers now — with a 50–75% smaller index alongside it.</p>
<h2 id="howtoenableit">How to enable it</h2>
<p>As of 9.0, <code>logs-*-*</code> data streams default to LogsDB automatically. If your data streams match that pattern, you're already using it.</p>
<blockquote>
  <p><strong>Want a hands-on walkthrough?</strong> <a href="https://www.elastic.co/observability-labs/blog/elasticsearch-logsdb-index-mode-storage-savings"><em>Cut Elasticsearch log storage costs by 76% with LogsDB</em></a> walks through creating two indices, reindexing, and measuring the difference with the <code>_stats</code> API — including version-specific enable instructions for 8.x clusters.</p>
</blockquote>
<p>For other index patterns, set it in your template:</p>
<pre><code>PUT _index_template/logs-template
{
  "index_patterns": ["logs-*"],
  "template": {
    "settings": {
      "index.mode": "logsdb"
    }
  }
}
</code></pre>
<p>Synthetic <code>_source</code> turns on automatically with <code>index.mode: logsdb</code>.</p>
<p>For the routing optimization (8.18+), add one more setting:</p>
<pre><code>PUT _index_template/logs-template
{
  "index_patterns": ["logs-*"],
  "template": {
    "settings": {
      "index.mode": "logsdb",
      "index.logsdb.route_on_sort_fields": true
    }
  }
}
</code></pre>
<p>This routes shards by sort field values instead of <code>_id</code>, adding ~20% storage reduction at a 1–4% ingestion penalty. It requires at least two sort fields beyond <code>@timestamp</code> and auto-generated <code>_id</code>.</p>
<p>Switching an existing index to LogsDB requires a reindex. So does rolling back. There's no in-place conversion, so try it on new data streams first.</p>
<p>Storage improves further as segments merge — freshly written data compresses well, but merged segments compress even better.</p>
<h2 id="whatsnext">What's next</h2>
<p>Elasticsearch still carries some structural overhead from its search engine roots. <code>_id</code> and <code>_seq_no</code> are two examples: both consume meaningful disk space (on small documents they can account for more than half the index size), but neither is essential for log analytics workloads.</p>
<p>We've already taken the first step for TSDB: <a href="https://github.com/elastic/elasticsearch/pull/144026">PR #144026</a> eliminated stored <code>_id</code> bytes from TSDB indices by reconstructing the field on the fly from doc values, the same approach synthetic <code>_source</code> uses. We're exploring the same direction for LogsDB.</p>
<p><strong>9.4 and beyond.</strong> The architecture still has room to improve, and we're on it.</p>
<p>For the full reference, see the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/current/logs-data-stream.html">logs data stream documentation</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/observability-labs/blog/elasticsearch-logsdb-storage-evolution</link>
    <guid isPermaLink="false">elasticsearch-logsdb-storage-evolution</guid>
    <category><![CDATA[Data Management]]></category>
    <category><![CDATA[Logs Analytics]]></category>
    <dc:creator><![CDATA[Luca Wintergerst]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt8dc9db2d94cde133/6a7f0a089090b01c7a84e8c5/elasticsearch-logsdb-storage-evolution.png" length="0" type="image/png"/>
    <pubDate>Thu, 09 Apr 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Connecting the Dots: ES|QL Joins for Richer Observability Insights]]></title>
    <description><![CDATA[Now in tech preview, ES|QL LOOKUP JOIN lets you enrich logs, metrics, and traces at query time no need to denormalize at ingest. Add deployment, infra, or business context dynamically, reduce storage, and accelerate root cause analysis in Elastic Obervability.]]></description>
    <content:encoded><![CDATA[<p>You might have seen our recent announcement about the <a href="https://www.elastic.co/blog/esql-lookup-join-elasticsearch">arrival of SQL-style joins in Elasticsearch</a> with ES|QL's LOOKUP JOIN command (now in Tech Preview!). While that post covered the basics, let's take a closer look at this in the context of Observability. How can this new join capability specifically help engineers and SREs make sense of their logs, metrics, and traces and make Elasticsearch more storage efficient by not denormalizing as much data?</p>
<p><strong>Note:</strong> Before we jump into the details, it’s important to mention again that this type of functionality today relies on a special lookup index. It is not (yet) possible to JOIN any arbitrary index.</p>
<p>Observability isn't just about collecting data; it's about understanding it. Often, the raw telemetry data – a log line, a metric point, a trace span – lacks the full context needed for quick diagnosis or impact assessment. We need to correlate data, enrich it with business or infrastructure context, and ask more advanced questions.</p>
<p>Historically, achieving this in Elasticsearch involved techniques like denormalizing data at ingest time (using ingest pipelines with enrich processors, for example) or performing joins client-side. </p>
<p>By adding the necessary context (like host details or user attributes) as data flowed in, each document arrived fully ready for queries and analytics without extra processing later on. This approach worked well in many cases and still does, particularly when the reference data changes slowly or when the enriched fields are critical for nearly every search. </p>
<p>However, as environments become more dynamic and diverse, the need to frequently update reference data (or avoid storing repetitive fields in every document) highlighted some of the trade-offs. </p>
<p>With the introduction of ES|QL LOOKUP JOIN in Elasticsearch 8.18 and 9.0, you now have an additional, more flexible option for situations where real-time lookups and minimal duplication are desired. Both methods—ingest-time enrichment and on-the-fly LOOKUP JOIN—complement each other and remain valid, depending on use case needs around update frequency, query performance, and storage considerations.</p>
<h2 id="whylookupjoinsforobservability">Why Lookup Joins for Observability</h2>
<p>Lookup joins keep things flexible. You can decide on the fly if you’d like to look up additional information to assist you in your investigation.</p>
<p>Here are some examples:</p>
<ul>
<li><p><strong>Deployment Information:</strong> Which version of the code is generating these errors?</p></li>
<li><p><strong>Infrastructure Mapping:</strong> Which Kubernetes cluster or cloud region is experiencing high latency? What hardware does it use?</p></li>
<li><p><strong>Business Context:</strong> Are critical customers being affected by this slowdown?</p></li>
<li><p><strong>Team Ownership:</strong> Which team owns the service throwing these exceptions?</p></li>
</ul>
<p>Keeping this kind of information perfectly denormalized onto <em>every single</em> log line or metric point can be challenging and inefficient. Lookup datasets – like lists of deployments, server inventories, customer tiers, or service ownership mappings – often change independently of the telemetry data itself.</p>
<p><code>LOOKUP JOIN</code> is ideal here because:</p>
<ol>
<li><p><strong>Lookup Indices are Writable:</strong> Update your deployment list, CMDB export, or on-call rotation in the lookup index, and your <em>next</em> ES|QL query immediately uses the fresh data. No need to re-run complex enrich policies or re-index data.</p></li>
<li><p><strong>Flexibility:</strong> You decide <em>at query time</em> which context to join. Maybe today you care about deployment versions, tomorrow about cloud regions.</p></li>
<li><p><strong>Simpler Setup:</strong> As the original post highlighted, there are no enrich policies to manage. Just create an index with <code>index.mode: lookup</code> and load your data - up to 2 billion documents per lookup index.</p></li>
</ol>
<h2 id="observabilityusecasesexampleswithesql">Observability Use Cases &amp; Examples with ES|QL</h2>
<p>Let’s now look at a few examples to see how Lookup Joins can help.</p>
<h3 id="enrichingerrorlogswithdeploymentcontext">Enriching Error Logs with Deployment Context</h3>
<p>Lets say you're seeing a spike in errors for your <code>checkout-service</code>. You have logs flowing into a data stream, but they only contain the service name. The documents don’t have any information about the deployment activity itself. </p>
<pre><code>FROM logs-*
&amp;nbsp; | WHERE log.level == "error"
&amp;nbsp;&amp;nbsp;| WHERE service.name == "opbeans-ruby"
</code></pre>
<p>You need to know if a recent deployment is contributing to these errors. To do this, we can maintain a <code>deployments_info_lkp</code> index (set with <code>index.mode: lookup</code>) that maps service names to their deployment times. This index could be updated from our CI/CD pipeline automatically any time a deployment happens.</p>
<pre><code>PUT /deployments_info_lkp
{
&amp;nbsp;&amp;nbsp;"settings": {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"index.mode": "lookup"
&amp;nbsp;&amp;nbsp;},
&amp;nbsp;&amp;nbsp;"mappings": {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"properties": {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"service": {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"properties": {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"name": {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"type": "keyword"
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;},
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"deployment_time": {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"type": "date"
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;},
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"version": {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"type": "keyword"
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;}
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;}
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;}
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;}
&amp;nbsp;&amp;nbsp;}
}
# Bulk index the deployment documents
POST /_bulk
{ "index" : { "_index" : "deployments_info_lkp" } }
{ "service.name": "opbeans-ruby", "service.version": "1.0", "deployment_time": "2025-05-22T06:00:00Z" }
{ "index" : { "_index" : "deployments_info_lkp" } }
{ "service.name": "opbeans-go", "service.version": "1.1.0", "deployment_time": "2025-05-22T06:00:00Z" }
</code></pre>
<p>Using this information you can now write a query that joins these two sources.</p>
<p><em>ES|QL Query:</em></p>
<pre><code>FROM logs-* 
&amp;nbsp; | WHERE log.level == "error"
&amp;nbsp;&amp;nbsp;| WHERE service.name == "opbeans-ruby"
&amp;nbsp;&amp;nbsp;| LOOKUP JOIN deployments_info_lkp ON service.name 
</code></pre>
<p>This alone is a good step towards troubleshooting the problem. You now have the deployment_time column available for each of your error documents. The last remaining step now is to use this for further filtering. </p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blte7de0d193189ba60/6a7f07b32f00b21366efe98a/discover.png" alt="Discover" /></p>
<p>Any of the data we managed to join from the lookup index can be handled as any other data we’d usually have available in the ES|QL query. This means that we can filter on it, and check if we had a recent deployment.</p>
<pre><code>FROM logs-*
  | WHERE log.level == "error"
  | WHERE service.name == "opbeans-ruby"
  | LOOKUP JOIN deployments_info_lkp ON service.name 
  | KEEP message, service.name, service.version, deployment_time 
  | WHERE deployment_time &gt; NOW() - 2h
</code></pre>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt906a78d7ecb20c61/6a7f07b642a117848495bcaa/discover2.png" alt="Discover2" /></p>
<h3 id="savingdiskspaceusingjoin">Saving disk space using JOIN</h3>
<p>Denormalizing data by including contextual information like host OS or cloud provider details directly in every log event is convenient for querying but can increase storage consumption, especially with high-volume data streams. Instead of storing this often-redundant information repeatedly, we can leverage joins to retrieve it on demand, potentially saving valuable disk space. While compression often handles repetitive data well, removing these fields entirely can still yield noticeable storage savings.</p>
<p>In this example we’ll use a dataset of 1,000,000 Kubernetes container logs using the default mapping of the Kubernetes integration, with <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/logs-data-stream">logsdb index mode</a> enabled. The starting size for this index is 35.5mb. </p>
<pre><code>GET _cat/indices/k8s-logs-default?h=index,pri.store.size
###&amp;nbsp;
k8s-logs-default &amp;nbsp; &amp;nbsp; &amp;nbsp; 35.5mb
</code></pre>
<p>Using the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-disk-usage">disk usage API</a>, we observed that fields like host.os and cloud.* contribute roughly 5% to the total index size on disk (35.5mb). These fields can be useful in some cases, but information like the os.name is rarely queried. </p>
<pre><code>// Example host.os structure
"os": {
&amp;nbsp;&amp;nbsp;"codename": "Plow", "family": "redhat", "kernel": "6.6.56+",
&amp;nbsp;&amp;nbsp;"name": "Red Hat Enterprise Linux", "platform": "rhel", "type": "linux", "version": "9.5 (Plow)"
}

// Example cloud structure
"cloud": {
&amp;nbsp;&amp;nbsp;"account": { "id": "elastic-observability" },
&amp;nbsp;&amp;nbsp;"availability_zone": "us-central1-c",
&amp;nbsp;&amp;nbsp;"instance": { "id": "5799032384800802653", "name": "gke-edge-oblt-edge-oblt-pool-46262cd0-w905" },
&amp;nbsp;&amp;nbsp;"machine": { "type": "e2-standard-4" },
&amp;nbsp;&amp;nbsp;"project": { "id": "elastic-observability" },
&amp;nbsp;&amp;nbsp;"provider": "gcp", "region": "us-central1", "service": { "name": "GCE" }
}
</code></pre>
<p>Instead of storing this information with every document, let's instead drop this information in an ingest pipeline.</p>
<pre><code>PUT _ingest/pipeline/drop-host-os-cloud
{
&amp;nbsp;&amp;nbsp;"processors": [
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;{ "remove": { "field": "host.os" } },
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;{ "set": { "field": "tmp1", "value": "{{cloud.instance.id}}" } }, // Temporarily store the ID
</code></pre>
<pre><code>&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;{ "remove": { "field": "cloud" } }, &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; // Remove the entire cloud object
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;{ "set": { "field": "cloud.instance.id", "value": "{{tmp1}}" } }, // Restore just the cloud instance ID
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;{ "remove": { "field": "tmp1", "ignore_missing": true } } &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; // Clean up temporary field
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;]
}
</code></pre>
<p>Reindexing (and <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-forcemerge">force merging to one segment</a>) now shows the following size, resulting in approximately 5% less space. </p>
<pre><code>GET _cat/indices/k8s-logs-*?h=index,pri.store.size
###&amp;nbsp;
k8s-logs-default &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; 33.7mb
k8s-logs-drop-cloud-os &amp;nbsp; &amp;nbsp; &amp;nbsp; 35.5mb
</code></pre>
<p>Now, to regain access to the removed host.os and cloud.* information during analysis without storing it in every log document, we can create a lookup index. This index will store the full host and cloud metadata, keyed by the cloud.instance.id that we preserved in our logs. This instance_metadata_lkp index will be significantly smaller than the space saved across millions or billions of log lines, as it only needs one document per unique instance.</p>
<pre><code># Create the lookup index for instance metadata
PUT /instance_metadata_lkp
{
&amp;nbsp;&amp;nbsp;"settings": {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"index.mode": "lookup"
&amp;nbsp;&amp;nbsp;},
&amp;nbsp;&amp;nbsp;"mappings": {
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"properties": {
</code></pre>
<pre><code>&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"cloud.instance.id": {&amp;nbsp; # The join key we kept in the logs
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"type": "keyword"
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;},
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"host.os": { &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; # The full host.os object we removed
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"type": "object",
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"enabled": false&amp;nbsp; &amp;nbsp; &amp;nbsp; # Often don't need to search sub-fields here
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;},
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"cloud": { &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; # The full cloud object we removed (mostly)
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"type": "object",
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;"enabled": false &amp;nbsp; &amp;nbsp; # Often don't need to search sub-fields here
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;}
&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;}
&amp;nbsp;&amp;nbsp;}
}

# Bulk index sample instance metadata (keyed by cloud.instance.id)
# This data might come from your cloud provider API or CMDB
POST /_bulk
{ "index" : { "_index" : "instance_metadata_lkp", "_id": "5799032384800802653" } }
{ "cloud.instance.id": "5799032384800802653", "host.os": { "codename": "Plow", "family": "redhat", "kernel": "6.6.56+", "name": "Red Hat Enterprise Linux", "platform": "rhel", "type": "linux", "version": "9.5 (Plow)" }, "cloud": { "account": { "id": "elastic-observability" }, "availability_zone": "us-central1-c", "instance": { "id": "5799032384800802653", "name": "gke-edge-oblt-edge-oblt-pool-46262cd0-w905" }, "machine": { "type": "e2-standard-4" }, "project": { "id": "elastic-observability" }, "provider": "gcp", "region": "us-central1", "service": { "name": "GCE" } } }
</code></pre>
<p>With this setup, when you need the full host or cloud context for your logs, you can simply use LOOKUP JOIN in your ES|QL query and continue filtering on the data from the lookup index</p>
<pre><code>FROM logs-*&amp;nbsp;
&amp;nbsp;&amp;nbsp;| LOOKUP JOIN instance_metadata_lkp ON cloud.instance.id 
&amp;nbsp; | WHERE cloud.region == "us-central1"
</code></pre>
<p>This approach allows us to query the full context when needed (e.g., filtering logs by host.os.name or cloud.region) while significantly reducing the storage footprint of the high-volume log indices by avoiding redundant data denormalization.</p>
<p>It should be noted that low cardinality metadata fields generally compress well and a large part of the storage savings in this case come from the “text” mapping of the host.os.name and cloud.instance.name field. Make sure to use the disk usage API to evaluate if this approach would be worth it in your specific use case. </p>
<h2 id="gettingstartedwithlookupsforobservability">Getting Started with Lookups for Observability</h2>
<p>Creating the necessary lookup indices is straightforward. As detailed in our <a href="http://link-to-original-blog-post">initial blog post</a>, you can use Kibana's Index Management UI, the Create Index API, or the File Upload utility – the key is setting <code>"index.mode": "lookup"</code> in the index settings.</p>
<p>For Observability, consider automating the population of these lookup indices:</p>
<ul>
<li><p>Export data periodically from your CMDB, CRM, or HR systems.</p></li>
<li><p>Have your CI/CD pipeline update the <code>deployments_lkp</code> index upon successful deployment.</p></li>
<li><p>Use tools like Logstash with an <code>elasticsearch</code> output configured to write to your lookup index.</p></li>
</ul>
<h2 id="anoteonperformanceandalternatives">A Note on Performance and Alternatives</h2>
<p>While incredibly powerful, joins aren't free. Each <code>LOOKUP JOIN</code> adds processing overhead to your query. For contextual data that is <em>very</em> static (e.g., the cloud region a host <em>permanently</em> resides in) and needed in <em>almost every</em> query against that data, the traditional approach of enriching at ingest time might still be slightly more performant for those specific queries, trading upfront processing and storage for query speed.</p>
<p>However, for the dynamic, flexible, and targeted enrichment scenarios common in Observability – like mapping to ever-changing deployments, user segments, or team structures – <code>LOOKUP JOIN</code> offers a compelling, efficient, and easier-to-manage solution.</p>
<h2 id="conclusion">Conclusion</h2>
<p>ES|QL's <code>LOOKUP JOIN</code> is making it easy to correlate and enrich your logs, metrics, and traces with up-to-date external information <em>at query time</em>; you can move faster from detecting problems to understanding their scope, impact, and root cause.</p>
<p>This feature is currently in Technical Preview in Elasticsearch 8.18 and Serverless, available now on Elastic Cloud. We encourage you to try it out with your own Observability data and share your feedback using the "Submit feedback" button in the ES|QL editor in Discover. We're excited to see how you use it to connect the dots in your systems!</p>]]></content:encoded>
    <link>https://www.elastic.co/observability-labs/blog/elastic-esql-join-observability</link>
    <guid isPermaLink="false">elastic-esql-join-observability</guid>
    <category><![CDATA[Logs Analytics]]></category>
    <dc:creator><![CDATA[Luca Wintergerst]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta7703fd6afaca645/6a7f07b93cab1c7ce20e4640/esql-join.jpg" length="0" type="image/jpeg"/>
    <pubDate>Thu, 29 May 2025 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Manual instrumentation of Go applications with OpenTelemetry]]></title>
    <description><![CDATA[In this blog post, we will show you how to manually instrument Go applications using OpenTelemetry. We will explore how to use the proper OpenTelemetry Go packages and, in particular, work on instrumenting tracing in a Go application.]]></description>
    <content:encoded><![CDATA[<p>DevOps and SRE teams are transforming the process of software development. While DevOps engineers focus on efficient software applications and service delivery, SRE teams are key to ensuring reliability, scalability, and performance. These teams must rely on a full-stack observability solution that allows them to manage and monitor systems and ensure issues are resolved before they impact the business.</p>
<p>Observability across the entire stack of modern distributed applications requires data collection, processing, and correlation often in the form of dashboards. Ingesting all system data requires installing agents across stacks, frameworks, and providers — a process that can be challenging and time-consuming for teams who have to deal with version changes, compatibility issues, and proprietary code that doesn't scale as systems change.</p>
<p>Thanks to <a href="http://opentelemetry.io">OpenTelemetry</a> (OTel), DevOps and SRE teams now have a standard way to collect and send data that doesn't rely on proprietary code and have a large support community reducing vendor lock-in.</p>
<p>In this blog post, we will show you how to manually instrument Go applications using OpenTelemetry. This approach is slightly more complex than using auto-instrumentation</p>
<p>In a <a href="https://www.elastic.co/blog/opentelemetry-observability">previous blog</a>, we also reviewed how to use the OpenTelemetry demo and connect it to Elastic<sup>®</sup>, as well as some of Elastic’s capabilities with OpenTelemetry. In this blog, we will use <a href="https://github.com/elastic/observability-examples">an alternative demo application</a>, which helps highlight manual instrumentation in a simple way.</p>
<p>Finally, we will discuss how Elastic supports mixed-mode applications, which run with Elastic and OpenTelemetry agents. The beauty of this is that there is <strong>no need for the otel-collector</strong>! This setup enables you to slowly and easily migrate an application to OTel with Elastic according to a timeline that best fits your business.</p>
<h2 id="applicationprerequisitesandconfig">Application, prerequisites, and config</h2>
<p>The application that we use for this blog is called <a href="https://github.com/elastic/observability-examples">Elastiflix</a>, a movie streaming application. It consists of several micro-services written in .NET, NodeJS, Go, and Python.</p>
<p>Before we instrument our sample application, we will first need to understand how Elastic can receive the telemetry data.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt13a3ccd8116bdc07/6a85ccb1f5f1a0cd052ec93b/GO-flowhcart.png" alt="Elastic configuration options for OpenTelemetry" /></p>
<p>All of Elastic Observability’s APM capabilities are available with OTel data. Some of these include:</p>
<ul>
<li>Service maps</li>
<li>Service details (latency, throughput, failed transactions)</li>
<li>Dependencies between services, distributed tracing</li>
<li>Transactions (traces)</li>
<li>Machine learning (ML) correlations</li>
<li>Log correlation</li>
</ul>
<p>In addition to Elastic’s APM and a unified view of the telemetry data, you will also be able to use Elastic’s powerful machine learning capabilities to reduce the analysis, and alerting to help reduce MTTR.</p>
<h2 id="prerequisites">Prerequisites</h2>
<ul>
<li>An Elastic Cloud account — <a href="https://cloud.elastic.co/">sign up now</a></li>
<li>A clone of the <a href="https://github.com/elastic/observability-examples">Elastiflix demo application</a>, or your own Go application</li>
<li>Basic understanding of Docker — potentially install <a href="https://www.docker.com/products/docker-desktop/">Docker Desktop</a></li>
<li>Basic understanding of Go</li>
</ul>
<h2 id="viewtheexamplesourcecode">View the example source code</h2>
<p>The full source code including the Dockerfile used in this blog can be found on <a href="https://github.com/elastic/observability-examples/tree/main/Elastiflix/go-favorite-otel-manual">GitHub</a>. The repository also contains the <a href="https://github.com/elastic/observability-examples/tree/main/Elastiflix/go-favorite">same application without instrumentation</a>. This allows you to compare each file and see the differences.</p>
<p>Before we begin, let’s look at the non-instrumented code first.</p>
<p>This is our simple go application that can receive a GET request. Note that the code shown here is a slightly abbreviated version.</p>
<pre><code>package main

import (
    "log"
    "net/http"
    "os"
    "time"

    "github.com/go-redis/redis/v8"

    "github.com/sirupsen/logrus"

    "github.com/gin-gonic/gin"
    "strconv"
    "math/rand"
)

var logger = &amp;logrus.Logger{
    Out:   os.Stderr,
    Hooks: make(logrus.LevelHooks),
    Level: logrus.InfoLevel,
    Formatter: &amp;logrus.JSONFormatter{
        FieldMap: logrus.FieldMap{
            logrus.FieldKeyTime:  "@timestamp",
            logrus.FieldKeyLevel: "log.level",
            logrus.FieldKeyMsg:   "message",
            logrus.FieldKeyFunc:  "function.name", // non-ECS
        },
        TimestampFormat: time.RFC3339Nano,
    },
}

func main() {
    delayTime,  := strconv.Atoi(os.Getenv("TOGGLE_SERVICE_DELAY"))

    redisHost := os.Getenv("REDIS_HOST")
    if redisHost == "" {
        redisHost = "localhost"
    }

    redisPort := os.Getenv("REDIS_PORT")
    if redisPort == "" {
        redisPort = "6379"
    }

    applicationPort := os.Getenv("APPLICATION_PORT")
    if applicationPort == "" {
        applicationPort = "5000"
    }

    // Initialize Redis client
    rdb := redis.NewClient(&amp;redis.Options{
        Addr:     redisHost + ":" + redisPort,
        Password: "",
        DB:       0,
    })

    // Initialize router
    r := gin.New()
    r.Use(logrusMiddleware)

    r.GET("/favorites", func(c *gin.Context) {
        // artificial sleep for delayTime
        time.Sleep(time.Duration(delayTime) * time.Millisecond)

        userID := c.Query("user_id")

        contextLogger(c).Infof("Getting favorites for user %q", userID)

        favorites, err := rdb.SMembers(c.Request.Context(), userID).Result()
        if err != nil {
            contextLogger(c).Error("Failed to get favorites for user %q", userID)
            c.String(http.StatusInternalServerError, "Failed to get favorites")
            return
        }

        contextLogger(c).Infof("User %q has favorites %q", userID, favorites)

        c.JSON(http.StatusOK, gin.H{
            "favorites": favorites,
        })
    })

    // Start server
    logger.Infof("App startup")
    log.Fatal(http.ListenAndServe(":"+applicationPort, r))
    logger.Infof("App stopped")
}
</code></pre>
<h2 id="stepbystepguide">Step-by-step guide</h2>
<h3 id="step0logintoyourelasticcloudaccount">Step 0. Log in to your Elastic Cloud account</h3>
<p>This blog assumes you have an Elastic Cloud account — if not, follow the <a href="https://cloud.elastic.co/registration?elektra=en-cloud-page">instructions to get started on Elastic Cloud</a>.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdec739045c430140/6a85ccb527c5cd10885f742e/elastic-blog-4-free-trial.png" alt="free trial" /></p>
<h3 id="step1installandinitializeopentelemetry">Step 1. Install and initialize OpenTelemetry</h3>
<p>As a first step, we’ll need to add some additional packages to our application.</p>
<pre><code>import (
      "github.com/go-redis/redis/extra/redisotel/v8"
      "go.opentelemetry.io/otel"
      "go.opentelemetry.io/otel/attribute"
      "go.opentelemetry.io/otel/exporters/otlp/otlptrace"
    "go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc"

    "go.opentelemetry.io/otel/propagation"

    "google.golang.org/grpc/credentials"
    "crypto/tls"

      sdktrace "go.opentelemetry.io/otel/sdk/trace"

    "go.opentelemetry.io/contrib/instrumentation/github.com/gin-gonic/gin/otelgin"

    "go.opentelemetry.io/otel/trace"
    "go.opentelemetry.io/otel/codes"
)
</code></pre>
<p>This code imports necessary OpenTelemetry packages, including those for tracing, exporting, and instrumenting specific libraries like Redis.</p>
<p>Next we read the "OTEL_EXPORTER_OTLP_ENDPOINT" variable and initialize the exporter.</p>
<pre><code>var (
    collectorURL = os.Getenv("OTEL_EXPORTER_OTLP_ENDPOINT")
)
var tracer trace.Tracer


func initTracer() func(context.Context) error {
    tracer = otel.Tracer("go-favorite-otel-manual")

    // remove https:// from the collector URL if it exists
    collectorURL = strings.Replace(collectorURL, "https://", "", 1)
    secretToken := os.Getenv("ELASTIC_APM_SECRET_TOKEN")
    if secretToken == "" {
        log.Fatal("ELASTIC_APM_SECRET_TOKEN is required")
    }

    secureOption := otlptracegrpc.WithInsecure()
    exporter, err := otlptrace.New(
        context.Background(),
        otlptracegrpc.NewClient(
            secureOption,
            otlptracegrpc.WithEndpoint(collectorURL),
            otlptracegrpc.WithHeaders(map[string]string{
                "Authorization": "Bearer " + secretToken,
            }),
            otlptracegrpc.WithTLSCredentials(credentials.NewTLS(&amp;tls.Config{})),
        ),
    )

    if err != nil {
        log.Fatal(err)
    }

    otel.SetTracerProvider(
        sdktrace.NewTracerProvider(
            sdktrace.WithSampler(sdktrace.AlwaysSample()),
            sdktrace.WithBatcher(exporter),
        ),
    )
    otel.SetTextMapPropagator(
        propagation.NewCompositeTextMapPropagator(
            propagation.Baggage{},
            propagation.TraceContext{},
        ),
    )
    return exporter.Shutdown
}
</code></pre>
<p>For instrumenting connections to Redis, we will add a tracing hook to it, and in order to instrument Gin, we will add the OTel middleware. This will automatically capture all interactions with our application, since Gin will be fully instrumented. In addition, all outgoing connections to Redis will also be instrumented.</p>
<pre><code>// Initialize Redis client
    rdb := redis.NewClient(&amp;redis.Options{
        Addr:     redisHost + ":" + redisPort,
        Password: "",
        DB:       0,
    })
    rdb.AddHook(redisotel.NewTracingHook())
    // Initialize router
    r := gin.New()
    r.Use(logrusMiddleware)
    r.Use(otelgin.Middleware("go-favorite-otel-manual"))
</code></pre>
<p><strong>Adding custom spans</strong><br />
Now that we have everything added and initialized, we can add custom spans.</p>
<p>If we want to have additional instrumentation for a part of our app, we simply start a custom span and then defer ending the span.</p>
<pre><code>// start otel span
ctx := c.Request.Context()
ctx, span := tracer.Start(ctx, "add_favorite_movies")
defer span.End()
</code></pre>
<p>For comparison, this is the instrumented code of our sample application. You can find the full source code in <a href="https://github.com/elastic/observability-examples/tree/main/Elastiflix/go-favorite-otel-manual">GitHub</a>.</p>
<pre><code>package main

import (
    "log"
    "net/http"
    "os"
    "time"
    "context"

    "github.com/go-redis/redis/v8"
    "github.com/go-redis/redis/extra/redisotel/v8"


    "github.com/sirupsen/logrus"

    "github.com/gin-gonic/gin"

  "go.opentelemetry.io/otel"
  "go.opentelemetry.io/otel/attribute"
  "go.opentelemetry.io/otel/exporters/otlp/otlptrace"
  "go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc"

    "go.opentelemetry.io/otel/propagation"

    "google.golang.org/grpc/credentials"
    "crypto/tls"

  sdktrace "go.opentelemetry.io/otel/sdk/trace"

    "go.opentelemetry.io/contrib/instrumentation/github.com/gin-gonic/gin/otelgin"

    "go.opentelemetry.io/otel/trace"

    "strings"
    "strconv"
    "math/rand"
    "go.opentelemetry.io/otel/codes"

)

var tracer trace.Tracer

func initTracer() func(context.Context) error {
    tracer = otel.Tracer("go-favorite-otel-manual")

    collectorURL = strings.Replace(collectorURL, "https://", "", 1)

    secureOption := otlptracegrpc.WithInsecure()

    // split otlpHeaders by comma and convert to map
    headers := make(map[string]string)
    for _, header := range strings.Split(otlpHeaders, ",") {
        headerParts := strings.Split(header, "=")

        if len(headerParts) == 2 {
            headers[headerParts[0]] = headerParts[1]
        }
    }

    exporter, err := otlptrace.New(
        context.Background(),
        otlptracegrpc.NewClient(
            secureOption,
            otlptracegrpc.WithEndpoint(collectorURL),
            otlptracegrpc.WithHeaders(headers),
            otlptracegrpc.WithTLSCredentials(credentials.NewTLS(&amp;tls.Config{})),
        ),
    )

    if err != nil {
        log.Fatal(err)
    }

    otel.SetTracerProvider(
        sdktrace.NewTracerProvider(
            sdktrace.WithSampler(sdktrace.AlwaysSample()),
            sdktrace.WithBatcher(exporter),
            //sdktrace.WithResource(resources),
        ),
    )
    otel.SetTextMapPropagator(
        propagation.NewCompositeTextMapPropagator(
            propagation.Baggage{},
            propagation.TraceContext{},
        ),
    )
    return exporter.Shutdown
}

var (
  collectorURL = os.Getenv("OTEL_EXPORTER_OTLP_ENDPOINT")
    otlpHeaders = os.Getenv("OTEL_EXPORTER_OTLP_HEADERS")
)


var logger = &amp;logrus.Logger{
    Out:   os.Stderr,
    Hooks: make(logrus.LevelHooks),
    Level: logrus.InfoLevel,
    Formatter: &amp;logrus.JSONFormatter{
        FieldMap: logrus.FieldMap{
            logrus.FieldKeyTime:  "@timestamp",
            logrus.FieldKeyLevel: "log.level",
            logrus.FieldKeyMsg:   "message",
            logrus.FieldKeyFunc:  "function.name", // non-ECS
        },
        TimestampFormat: time.RFC3339Nano,
    },
}

func main() {
    cleanup := initTracer()
  defer cleanup(context.Background())

    redisHost := os.Getenv("REDIS_HOST")
    if redisHost == "" {
        redisHost = "localhost"
    }

    redisPort := os.Getenv("REDIS_PORT")
    if redisPort == "" {
        redisPort = "6379"
    }

    applicationPort := os.Getenv("APPLICATION_PORT")
    if applicationPort == "" {
        applicationPort = "5000"
    }

    // Initialize Redis client
    rdb := redis.NewClient(&amp;redis.Options{
        Addr:     redisHost + ":" + redisPort,
        Password: "",
        DB:       0,
    })
    rdb.AddHook(redisotel.NewTracingHook())


    // Initialize router
    r := gin.New()
    r.Use(logrusMiddleware)
    r.Use(otelgin.Middleware("go-favorite-otel-manual"))


    // Define routes
    r.GET("/", func(c *gin.Context) {
        contextLogger(c).Infof("Main request successful")
        c.String(http.StatusOK, "Hello World!")
    })

    r.GET("/favorites", func(c *gin.Context) {
        // artificial sleep for delayTime
        time.Sleep(time.Duration(delayTime) * time.Millisecond)

        userID := c.Query("user_id")

        contextLogger(c).Infof("Getting favorites for user %q", userID)

        favorites, err := rdb.SMembers(c.Request.Context(), userID).Result()
        if err != nil {
            contextLogger(c).Error("Failed to get favorites for user %q", userID)
            c.String(http.StatusInternalServerError, "Failed to get favorites")
            return
        }

        contextLogger(c).Infof("User %q has favorites %q", userID, favorites)

        c.JSON(http.StatusOK, gin.H{
            "favorites": favorites,
        })
    })

    // Start server
    logger.Infof("App startup")
    log.Fatal(http.ListenAndServe(":"+applicationPort, r))
    logger.Infof("App stopped")
}
</code></pre>
<h3 id="step2runningthedockerimagewithenvironmentvariables">Step 2. Running the Docker image with environment variables</h3>
<p>As specified in the <a href="https://opentelemetry.io/docs/specs/otel/configuration/sdk-environment-variables/">OTEL documentation</a>, we will use environment variables and pass in the configuration values that are found in your APM Agent’s configuration section.</p>
<p>Because Elastic accepts OTLP natively, we just need to provide the Endpoint and authentication where the OTEL Exporter needs to send the data, as well as some other environment variables.</p>
<p><strong>Where to get these variables in Elastic Cloud and Kibana</strong> <sup>®</sup><br />
You can copy the endpoints and token from Kibana under the path /app/home#/tutorial/apm.</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltebfc9096105dc59e/6a85ccb89d2b7100f0f939ca/elastic-blog-GO-apm-agents.png" alt="GO apm agents" /></p>
<p>You will need to copy the OTEL_EXPORTER_OTLP_ENDPOINT as well as the OTEL_EXPORTER_OTLP_HEADERS.</p>
<p><strong>Build the image</strong></p>
<pre><code>docker build -t  go-otel-manual-image .
</code></pre>
<h2 id="runtheimage">Run the image</h2>
<pre><code>docker run \
       -e OTEL_EXPORTER_OTLP_ENDPOINT="&lt;REPLACE WITH OTEL_EXPORTER_OTLP_ENDPOINT&gt;" \
       -e OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer &lt;REPLACE WITH TOKEN&gt;" \
       -e OTEL_RESOURCE_ATTRIBUTES="service.version=1.0,deployment.environment=production,service.name=go-favorite-otel-manual" \
       -p 5000:5000 \
       go-otel-manual-image
</code></pre>
<p>You can now issue a few requests in order to generate trace data. Note that these requests are expected to return an error, as this service relies on a connection to Redis that you don’t currently have running. As mentioned before, you can find a more complete example using Docker compose <a href="https://github.com/elastic/observability-examples/tree/main/Elastiflix">here</a>.</p>
<pre><code>curl localhost:500/favorites
# or alternatively issue a request every second

while true; do curl "localhost:5000/favorites"; sleep 1; done;
</code></pre>
<h2 id="howdothetracesshowupinelastic">How do the traces show up in Elastic?</h2>
<p>Now that the service is instrumented, you should see the following output in Elastic APM when looking at the transactions section of your Node.js service:</p>
<p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blta64d1c7fc7f17c32/6a85ccbb2d64d5249d081d72/GO-trace-samples.png" alt="trace samples" /></p>
<h2 id="conclusion">Conclusion</h2>
<p>In this blog, we discussed the following:</p>
<ul>
<li>How to manually instrument Go with OpenTelemetry</li>
<li>How to properly initialize OpenTelemetry and add a custom span</li>
<li>How to easily set the OTLP ENDPOINT and OTLP HEADERS with Elastic without the need for a collector</li>
</ul>
<p>Hopefully, this provides an easy-to-understand walk-through of instrumenting Go with OpenTelemetry and how easy it is to send traces into Elastic.</p>
<blockquote>
  <p>Developer resources:</p>
  <ul>
  <li><a href="https://www.elastic.co/blog/getting-started-opentelemetry-instrumentation-sample-app">Elastiflix application</a>, a guide to instrument different languages with OpenTelemetry</li>
  <li>Python: <a href="https://www.elastic.co/blog/auto-instrumentation-of-python-applications-opentelemetry">Auto-instrumentation</a>, <a href="https://www.elastic.co/blog/manual-instrumentation-of-python-applications-opentelemetry">Manual-instrumentation</a></li>
  <li>Java: <a href="https://www.elastic.co/blog/auto-instrumentation-of-java-applications-opentelemetry">Auto-instrumentation</a>, <a href="https://www.elastic.co/blog/manual-instrumentation-of-java-applications-opentelemetry">Manual-instrumentation</a></li>
  <li>Node.js: <a href="https://www.elastic.co/blog/auto-instrument-nodejs-applications-opentelemetry">Auto-instrumentation</a>, <a href="https://www.elastic.co/blog/manual-instrumentation-of-nodejs-applications-opentelemetry">Manual-instrumentation</a></li>
  <li>.NET: <a href="https://www.elastic.co/blog/auto-instrumentation-of-net-applications-opentelemetry">Auto-instrumentation</a>, <a href="https://www.elastic.co/blog/manual-instrumentation-of-net-applications-opentelemetry">Manual-instrumentation</a></li>
  <li>Go: <a href="https://elastic.co/blog/manual-instrumentation-apps-opentelemetry">Manual-instrumentation</a></li>
  <li><a href="https://www.elastic.co/blog/best-practices-instrumenting-opentelemetry">Best practices for instrumenting OpenTelemetry</a></li>
  </ul>
  <p>General configuration and use case resources:</p>
  <ul>
  <li><a href="https://www.elastic.co/blog/opentelemetry-observability">Independence with OpenTelemetry on Elastic</a></li>
  <li><a href="https://www.elastic.co/blog/implementing-kubernetes-observability-security-opentelemetry">Modern observability and security on Kubernetes with Elastic and OpenTelemetry</a></li>
  <li><a href="https://www.elastic.co/blog/3-models-logging-opentelemetry-elastic">3 models for logging with OpenTelemetry and Elastic</a></li>
  <li><a href="https://www.elastic.co/blog/adding-free-and-open-elastic-apm-as-part-of-your-elastic-observability-deployment">Adding free and open Elastic APM as part of your Elastic Observability deployment</a></li>
  <li><a href="https://www.elastic.co/blog/custom-metrics-app-code-java-agent-plugin">Capturing custom metrics through OpenTelemetry API in code with Elastic</a></li>
  <li><a href="https://www.elastic.co/virtual-events/future-proof-your-observability-platform-with-opentelemetry-and-elastic">Future-proof your observability platform with OpenTelemetry and Elastic</a></li>
  <li><a href="https://www.elastic.co/blog/kubernetes-k8s-observability-elasticsearch-cncf">Elastic Observability: Built for open technologies like Kubernetes, OpenTelemetry, Prometheus, Istio, and more</a></li>
  </ul>
</blockquote>
<p>Don’t have an Elastic Cloud account yet? <a href="https://cloud.elastic.co/registration">Sign up for Elastic Cloud</a> and try out the auto-instrumentation capabilities that I discussed above. I would be interested in getting your feedback about your experience in gaining visibility into your application stack with Elastic.</p>
<p><em>The release and timing of any features or functionality described in this post remain at Elastic's sole discretion. Any features or functionality not currently available may not be delivered on time or at all.</em></p>]]></content:encoded>
    <link>https://www.elastic.co/observability-labs/blog/manual-instrumentation-apps-opentelemetry</link>
    <guid isPermaLink="false">manual-instrumentation-apps-opentelemetry</guid>
    <category><![CDATA[OpenTelemetry]]></category>
    <category><![CDATA[APM]]></category>
    <category><![CDATA[Infrastructure Monitoring]]></category>
    <dc:creator><![CDATA[Luca Wintergerst]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt985f77895b54aaab/6a85ccbe342d69d08921b121/observability-launch-series-5-go-manual.jpg" length="0" type="image/jpeg"/>
    <pubDate>Tue, 12 Sep 2023 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>