Blog

Kubernetes noisy neighbor in a multi-tenant cluster: Find it with ES|QL

Every pod was Running, and storefront requests under 500 ms still fell from 99.6% to 59%. ES|QL on kubelet metrics found a batch job taking 1.7 of the node's 2 cores, and a 500m cap brought the storefront back above 99%.

A Kubernetes noisy neighbor cut the share of storefront requests served in under 500 ms from 99.6% to 59%. Every pod on the node stayed running the whole time, with no restarts or alerts. One batch workload was using 1.7 of the worker's two cores, leaving the storefront with 60% of the CPU it needed. Elasticsearch Query Language (ES|QL) queries on kubelet metrics from the Elastic Distribution of OpenTelemetry (EDOT) Collector find it and trace it to its deployment, and capping that deployment at 500m brings 99.8% of requests back under 500 ms. A ResourceQuota holds the namespace to that budget, and a latency Service Level Objective (SLO) tracks the storefront from then on.

Prerequisites

  • Elasticsearch and Kibana 9.4+. The TS command.

  • A Kubernetes cluster with the EDOT Collector installed through the Kubernetes quickstart, plus kubectl and Helm.

  • A latency metric for the affected service. The example uses a k6 probe that exports k6_http_req_duration through OpenTelemetry (OTel).

  • A Kibana API key with SLO write privileges for the last section.

You can find the companion notebook here. It reproduces the whole lab from this article, including the tenant labels and the SLO, and cleans up at the end.

The scenario: A batch export and a storefront in a multi-tenant cluster

Two tenants share the worker node of a two-node kind cluster. The worker is limited to two CPUs.

  • team-storefront runs an NGINX deployment that gzip-compresses an 8 MiB response. It requests 10m of CPU and has no CPU limit.

  • team-batch runs batch-export, a stress container with four CPU workers and 384 MiB of allocated memory. It requests one CPU with a limit of two.

  • A k6 probe outside the worker sends one request at a time with a 200 ms pause. A good request returns HTTP 200 in under 500 ms. Running alone, the storefront answers 99.6% of the probe's requests within that threshold.

The worker's MemoryPressure condition stays false for the whole run, so this is a CPU contention case. The Docker limit is a CFS quota that the kubelet doesn’t see, so Kubernetes still reports the Docker virtual machine’s (VM's) six CPUs as the node's capacity, and the queries compare CPU cores consumed instead of a percentage of capacity.

Step 1: Add tenant labels to Kubernetes metrics

Put the ownership labels on each deployment's pod template, where the collector can read them from the pod:

How helpful was this content?

Related Content

Signed, sealed, delivered: ECK certificate management with Vault and cert-manager

Signed, sealed, delivered: ECK certificate management with Vault and cert-manager

Anand Vyas
Elastic Cloud on Kubernetes, simplified: zone awareness, restarts, and mTLS

Elastic Cloud on Kubernetes, simplified: zone awareness, restarts, and mTLS

Omer Kushmaro
Migrating your OpenShift Elasticsearch 6.x cluster to Elastic Cloud on Kubernetes (ECK)

Migrating your OpenShift Elasticsearch 6.x cluster to Elastic Cloud on Kubernetes (ECK)

Omer Kushmaro
AutoOps in action: Investigating Elasticsearch cluster performance on ECK

AutoOps in action: Investigating Elasticsearch cluster performance on ECK

Aram Favela
Multi-tenancy in Elastic Cloud on Kubernetes deployments: Example architectures

Multi-tenancy in Elastic Cloud on Kubernetes deployments: Example architectures

Lorenzo Soligo

Elastic Observability Labs Newsletter