<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Mappings - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Mappings - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/search-labs/blog/category/mappings</link>
    </image>
    <link>https://www.elastic.co/search-labs/blog/category/mappings</link>
    <atom:link href="https://www.elastic.co/search-labs/rss/category/mappings.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Wed, 07 Oct 2026 17:49:42 GMT</lastBuildDate>
  <item>
    <title><![CDATA[One field, one copy: How Elasticsearch columnar storage drops the inverted index]]></title>
    <description><![CDATA[Storing each field once means no inverted index, so doc values now read in bulk and skippers let queries skip whole ranges of documents, while new mapping attributes control what each field is allowed to contain.]]></description>
    <content:encoded><![CDATA[<p>As part of the 9.5.0 release, Elasticsearch introduced <a href="https://www.elastic.co/docs/reference/elasticsearch/columnar">columnar and logsdb_columnar index modes</a> in technical preview. Elasticsearch has had columnar storage using Lucene’s doc values since version 1.0.0. Lucene’s doc values power analytics and search functionalities, like group by and sorting by a field. So, what changes with the columnar index modes? </p><p>The changes are about storage and performance, along with the out-of-the-box (OOTB) experience. Up until 9.5.0, Elasticsearch operated as a document-based search engine by default. It could be set up to behave like a columnar system storage-wise, but that wasn’t the OOTB experience. <a href="https://www.elastic.co/docs/reference/elasticsearch/columnar">Columnar index modes</a> make a number of fundamental changes that allow Elasticsearch to optimize columnar analytic and search use cases:</p><ul><li><p>Fields are stored once as doc values only and are no longer indexed by default.</p></li><li><p>New multi-value semantics. The original ordering of multiple values per field per document (for example, in arrays) is preserved by default.</p></li><li><p>Mappings are always flat, and object and passthrough fields in mappings are always auto-flattened.</p></li></ul><h2>How columnar index modes fit into Elasticsearch</h2><p>Many of the columnar index mode changes originate from <a href="https://www.elastic.co/docs/manage-data/data-store/data-streams/time-series-data-stream-tsds">time series data streams (TSDSs)</a>. As part of making TSDB a competitive metrics solution, we improved doc values format on disk and only store dimensions and metric fields once as doc values. We also improved query performance. TSDB is already columnar today. Essentially, this makes TSDB’s storage mode columnar. The lessons learned from TSDB are now being applied more broadly to Elasticsearch. </p><p>Note that columnar index modes are opt-in and columnar, and document-based indices can coexist in the same cluster. An enterprise search use case can use a document-oriented index mode, while a logging use case can use logsdb_columnar index mode and be fully columnar, all in the same cluster. In fact, there are currently seven index modes, and indices can all use them in the same cluster.</p><h2>How columnar storage stays fast without an inverted index</h2><p>Indexed fields, either an inverted index for string-based fields or block k-dimensional (BKD) tree for numeric fields, allow Elasticsearch to query or filter by field very efficiently. However, the cost for this is an additional expensive data structure that uses a lot of disk space and is expensive to build at index time and at merge time. With the columnar index modes, fields are no longer indexed by default, so what did we do for query performance to be still acceptable on fields that were no longer indexed?</p><p>One major change was improving doc values scanning performance. This is key and is the cornerstone that any columnar system relies on. Previously, the scanning of doc values was essentially document by document. Lucene’s doc values API only allowed for looking up one value at a time. Historically, this fit the execution model of a search engine. In our own doc value format, we build the capability to allow bulk reading of values for Elasticsearch Query Language (ES|QL) queries. Also over recent minor Lucene releases, Lucene doc values API added support for bulk reading. Without this, fast columnar scanning wouldn’t have been possible.</p><p>Secondly, we fully adopted <a href="https://www.elastic.co/search-labs/blog/docvaluesskippers-lucene-range-queries">doc value skippers</a>, a hierarchical skiplist over doc values. Contrary to an inverted index or a BKD tree, doc values skippers are lightweight data structures. At its core, a <em>skipper</em> allows queries to skip over a range of documents that don’t match a query. It can do this because it stores information, like min and max values. So, for example, when a range query is executed, an interval of documents can be skipped based on the intermediate result and a doc value skipper’s min and max values. The effectiveness of doc value skippers depends on the order in which documents are laid out on disk. This is why <a href="https://www.elastic.co/docs/reference/elasticsearch/columnar#index-sorting">index sorting</a> should be enabled or altered to match the use case.</p><p>By significantly improving our columnar scanning and doubling down on doc values skippers, we’re able to avoid indexing fields by default. Note that a field can still be indexed; <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/mapping-index">the <code>index</code> mapping attribute</a> just defaults to <code>false</code> in columnar mode. The exceptions to this rule are text-based fields, which are still indexed by default. This is because text fields provide free text search, which includes text analysis, along with phrase and wildcard matching. This is different from just filtering.</p><h2>How Elasticsearch handles high and low cardinality fields</h2><p>When setting up a schema with string fields, an important configuration parameter is often <em>cardinality</em>; that is, whether many unique values or a few unique string values are expected. Many systems have dedicated field or column types that target low and high cardinality string fields.</p><p>Fields that have low cardinality are typically stored with a dictionary, containing all unique values. Then, for each row offset, the offset into the dictionary containing the term the row has is stored. This is often called an <em>ordinal</em>. For low cardinality fields, this works well, as storing an ordinal per document takes up much less space. Encoding techniques, like delta encoding, offset encoding, and bitpacking, work well for ordinals to compact the per-document storage to just a few bits. </p><p>However, for high cardinality fields, the dictionary and ordinal approach can work counterintuitively. If a larger percentage of the documents have a unique value, building the dictionary becomes expensive and storage savings diminish. The dictionary then becomes another level of indirection for reading values. This is why most systems in that case store values in a columnar fashion using block-based compression. For example, values of multiple rows are stored in 128KB blocks using a sliding-window dictionary-based compression algorithm (like zstandard or lz4). This, in general, is a simple and effective method to store higher cardinality fields and avoids building and maintaining a dictionary. </p><p>With document-based Elasticsearch, there are two ways to map a string: using either the keyword field mapping or one of the text-based field mappings. The former stores an inverted index and dictionary-based doc values. The latter only stores an inverted index. This is why, typically, a text field mapper is often used in combination with a keyword mapper as a <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/multi-fields">multi-field</a>. Also, keyword field mapper (as the name suggests) is meant for keywords or fields that have a lower cardinality and uses dictionary-based doc values implementation. However, in practice, keyword field mappers are also used for high cardinality.</p><p>In columnar mode, every field is only stored once by default. For keyword fields, this means only doc values are stored with no inverted index. Text-based fields now also store doc values and an inverted index by default. Text-based field mappers are different from the keyword field mapper, as these are not automatically used via Elasticsearch’s dynamic mapping logic for columnar indices, and therefore text fields keep storing an inverted index by default.</p><p>For both keyword- and text-based fields, we didn’t choose to expose a cardinality mapping attribute. It’s not always possible to know ahead of time whether a field is low or high cardinality. When flushing and merging segments to disk, Elasticsearch sees all values and can determine the cardinality of a field. This is why we’re choosing to automatically determine whether the usage of a dictionary and ordinal-based encoding is beneficial over block-based compression using a simple cardinality threshold. If a field is below this threshold, dictionary and ordinal-based encoding is used; otherwise block-based compression is used. This simplifies configuration of the mappings and makes it possible to automatically optimize storage as data evolves, since some segments may use dictionaries while others may use blocks for the same field. However, this is currently not ready yet and so, as part of 9.5.0, in columnar mode, both keyword- and text-based field mappers store values in doc values in a block-based compressed layout on disk.</p><p>The two approaches compare as follows:</p><p>
</p><p><strong>Dictionary and ordinal encoding</strong></p><p><strong>Block-based compression</strong></p><p>Suits</p><p>Low cardinality fields</p><p>High cardinality fields</p><p>What’s stored</p><p>A dictionary of unique values, plus one ordinal per document</p><p>Values for many documents compressed together in blocks</p><p>Compression</p><p>Delta encoding, offset encoding, and bitpacking reduce each ordinal to a few bits</p><p>Sliding-window dictionary compression, such as zstandard or lz4, typically over 128KB blocks</p><p>Read path</p><p>Resolve the ordinal, and then look up the value in the dictionary</p><p>Decompress the block, and then read the value directly</p><p>Cost as cardinality rises</p><p>Dictionary grows large, savings shrink, and the extra indirection stays</p><p>Stable, with no dictionary to build or maintain</p><p>Used in columnar mode tech preview</p><p>Not yet</p><p>Yes, for both keyword and text fields</p><h2>Columnar mapping attributes: multi_value, nullability, on_failure</h2><p>The columnar index modes provide more control over how data is stored as doc values. By default, Elasticsearch is lenient and accepts all non-malformed values (for example, nulls and multiple values per field and document). If documents have fields with multiple values per document or no value, doc values store additional data structures to deal with them and therefore implicitly increase costs. </p><p>With columnar, new mapping attributes provide additional control. Note that these new mapping attributes are currently only available with the columnar index modes but will eventually also be available for all index modes.</p><p>Three new mapping attributes are involved:</p><p><strong>Attribute</strong></p><p><strong>Default</strong></p><p><strong>Enforces</strong></p><p><strong>On violation</strong></p><p><strong>Available</strong></p><p><code>multi_value</code></p><p><code>true</code></p><p>One value per document per field</p><p>Document indexing fails</p><p>9.5.0</p><p><code>nullability</code></p><p><code>true</code></p><p>Field must have a value</p><p>Document indexing fails</p><p>9.5.0</p><p><code>on_failure</code></p><p><code>fail</code></p><p>How the above failures are handled</p><p>Sets <code>fail</code> or <code>ignore</code> behavior</p><p>Next minor release</p><h3>multi_value: Enforcing single-valued fields</h3><p>By default, Elasticsearch accepts multiple values per document. To understand the implications of this, we first should take a look at how Elasticsearch (using Lucene’s doc values) stores a dense numeric field where all documents have a single value: </p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt3abb1f6fb702b165/6a9fa698893681d9ef3df7d6/unnamed.png" alt="Columnar storage doc values layout: value blocks and a block index resolve which value belongs to a docId" /><p>With this layout, all values are stored in blocks. The number of values per block depends on the index mode but is typically 128 values and is always the same within an index. All values in a block are encoded using various encoding techniques, like delta encoding and bit packing, so each block can have a different size, depending on how well the encoding techniques compress the values. This is why a block index is required. </p><p>Lucene has the notion of a docid (internal numbering for a document), which is essentially a row identifier.</p><ol><li><p>Queries produce matching docids.</p></li><li><p>In case of a dense field, the block id can be resolved from the docid directly.</p></li><li><p>The offset of a block can be resolved from the block index.</p></li><li><p>Once that has been looked up, the target block gets decoded and all values are available.</p></li><li><p>Finally, from docid, the ordinal within the decoded values array can be resolved, which produces the final value.</p></li></ol><p>Now let’s have a look at how the data layout changes when documents have multiple values per document:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9d5ca280a19c7a70/6a9fa6edee57e5dec9051db9/unnamed.png" alt="Multi-value doc values layout in columnar storage: offsets map a docId to values spanning value blocks" /><p>To determine how many values belong to a single docid, an offset lookup is required.</p><p>In this case, a docid has one or more offsets. Each offset points to a block index. Values for a single document are adjacent but can stretch over blocks. In general, compaction of values works well in blocks because values are similar. However, multi-value fields can cause the compaction of values to be less efficient if the number of values per field and document is large and values aren’t similar. This and the additional storage of offsets result in multi-value fields typically having a higher storage footprint on disk.</p><p></p><p>If a field is truly single-valued, you may want to enforce this property. <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/doc-values#doc-values-multi-value">The <code>multi_value</code> mapping attribute</a> makes that possible now. An example is a log level field. Logs typically have one log level (such as debug, info, or error). Enforcing that this field is single-valued in your mappings can help avoid accidentally using more storage than anticipated. A mapping snippet example that disallows the field <code>log.level</code> to have multiple values per document:</p>{
	"properties": {
		"log.level": {
			"type": "keyword",
			"multi_value": false
		}
	}
}<p></p><p>Note that even if a field allows multiple values, this doesn’t mean an offset lookup is stored. This only happens when a Lucene segment has at least one document with two or more values. The <code>multi_value</code> mapping attribute exists just for enforcement.</p><h3>nullability: Requiring every document to have a value</h3><p>By default, Elasticsearch accepts documents with fields that have no value or null value. Just as  multiple values per document require additional accounting, documents with no value require additional accounting to identify which of them have at least one value.</p><p>Doc values store a docid to offset lookup (known as IndexedDISI in Lucene) in case not all documents have a value in a segment. The offset either points directly to the block index for single-valued fields or the offset lookup in case of multi-valued fields. This lookup is compact compared to the value blocks being stored. However, if it were to be created for fields that should have at least one value per document, that would be a waste.</p><p><a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/doc-values#doc-values-nullability">The <code>nullability</code> mapping attribute</a> allows you to control whether documents are allowed to have no value. Just like the <code>multi_value</code> mapping attribute, the <code>nullability</code> attribute exists for enforcement.  Following is a mapping snippet example that requires the <code>log.level</code> field to have a value:</p>{
  "properties": {
     "log.level": {
        "type": "keyword",
        "nullability": false
     }
  }
}<h3>on_failure: What happens when validation fails</h3><p>What happens if a document has multiple values for a field and if the <code>multi_value</code> mapping attribute is set to <code>false</code> or when a field is mapped with nullability set to <code>false</code> and a document doesn’t have that field? At the moment, indexing such documents will fail with a bad request error.</p><p>As part of the next minor release, <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/doc-values#doc-values-on-failure">the <code>on_failure</code> mapping attribute</a> will be available. This allows you to indicate how to handle these validation failures, on a per-mapped field basis. This will support two values:</p><ol><li><p>Fail: Fail indexing of the entire document with a client error. This is the current behavior in Elasticsearch 9.5.0.</p></li><li><p>Ignore: Ignore the validation error, mark the field as ignored, and store values for that field in a hidden field so that it can be introspected when requesting the source. </p></li></ol><h2>Trying the columnar index modes</h2><p>The <a href="https://www.elastic.co/docs/reference/elasticsearch/columnar">columnar index modes</a> are still under active development, but we encourage you to give them a test drive. As we prepare the columnar index modes for general availability (GA), we’ll add more performance and efficiency improvements. We believe that by adapting a columnar mindset, many use cases will benefit from being more cost effective or having better performance characteristics.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/columnar-storage-elasticsearch-index-modes</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/columnar-storage-elasticsearch-index-modes</guid>
    <category><![CDATA[Lucene]]></category>
    <category><![CDATA[Index Data]]></category>
    <category><![CDATA[Mappings]]></category>
    <dc:creator><![CDATA[Martijn van Groningen]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltef1303bd86788300/6a9fa5cdf08ee10715855390/unnamed.png" length="0" type="image/png"/>
    <pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Taming PUNKs: How ES|QL queries Elasticsearch fields it was never told about]]></title>
    <description><![CDATA[In Elasticsearch 9.5, ES|QL can query unmapped fields. It reads them from _source or returns nulls, so a query keeps working when a field drops out of the mapping and you avoid a reindex that takes hours.]]></description>
    <content:encoded><![CDATA[<p>How do you make an analytical query engine use data that it cannot know exists? You “just” read the query, since everything that the user asks for is right there. Right?</p><p>In Elasticsearch 9.5, <a href="https://www.elastic.co/docs/reference/query-languages/esql">Elasticsearch Query Language (ES|QL)</a> queries no longer fail when a field isn't in the mapping. The new <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-unmapped-fields"><code>unmapped_fields</code> setting</a> lets queries load values from <code>_source</code> or fill with <code>nulls</code>, so queries keep working even when a backing index changes and a field goes missing, and can use unmapped data without reindexing. Here’s how we built that: the design choices and the edge cases (including a class of fields we nicknamed PUNKs), along with the testing strategies that gave us the confidence to ship it in general availability (GA).</p><h2>Why ES|QL queries fail when a field is unmapped</h2><p>You built a visualization using an ES|QL query. You refined it, and the query grew. You’re at 15 chained commands and counting, but it does <em>just</em> the right thing. It works, and your dashboard is <em>useful</em>.</p><p>Your query uses an index from a remote cluster, say <code>my-remote:logs-foo</code>. But actually, <code>logs-foo</code> is an alias, and at some point, the remote cluster makes it point to a different backing index. The new index is missing a field that’s used in your query, and your query and visualization break.</p><p>Or maybe you have an already fairly large index, and while building ES|QL queries on top of it, you realize that you’d like to use a field in the indexed documents that unfortunately never made it into the index mapping. You could reindex the data, but that would take hours.</p><p>ES|QL’s <code>unmapped_fields</code> setting is meant to deal with these types of situations.</p><p>If your query looks like this:</p><p>and <code>some_field</code> is unmapped, ES|QL’s default behavior is to fail with a verification exception.</p><p>You can use the <code>unmapped_fields</code> setting to instead either fill <code>some_field</code> with <code>null</code>s or read it from the document’s <code>_source</code>, like so:</p><h2>How ES|QL resolves queries with field caps</h2><p>Before we jump into the inner workings of <code>unmapped_fields</code>, we have to look into how ES|QL resolves queries regularly. Let’s consider the above query:</p><p>We said that if <code>some_field</code> isn’t in the mapping for <code>index</code>, ES|QL will reject the query. How does it make that decision?</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt7ba9af4b4b550018/6a8ee83fe41d7fea88654d42/image4.png" alt="ES|QL query resolution flow: analyzer checks index mappings, unresolved fields fail with Unknown column error" /><h3>How field caps tells ES|QL which fields exist</h3><p>In a typical schema-on-write fashion, Elasticsearch clusters maintain mappings with their respective indices. As a first step, ES|QL makes an internal request to the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-field-caps">field caps endpoint</a> to determine which fields the <code>index</code> has. It then passes the query, together with the field caps response, to the query planner, which consists essentially of the query analyzer (unrelated to analyzers of text fields) and query optimizer. The analyzer makes sense of raw names, like <code>some_field</code>, and notices that they correspond to index fields (or not). If all went well, the query is then passed on to the optimizer, which rewrites the query for efficiency, before it’s handed to the compute engine for execution.</p><h3>How the analyzer resolves field names in the query plan</h3><p>Let’s zoom in to the analyzer. The parsed query is represented in a tree structure, and the analyzer partially rewrites it, one command at a time, until it either has resolved all references or not.</p><p>For illustration, let’s use a somewhat more complex query and see how the analyzer would resolve it:</p><p>The parsed tree is actually a chain here, and it looks something like this:</p><p></p><p>The analyzer then moves up through the query tree to try and resolve the field names used in every command.</p><p>This is a simplified version of how we represent parse trees in tests and when debugging. The bottom of the chain corresponds to the <code>FROM</code> command and contains a list of all mapped fields that we know about, obtained from the field caps endpoint. (The <code>{f}</code> suffix marks an actually mapped field for better distinction later.)</p><p>The two <code>EVAL</code> nodes on top of it correspond to the remaining commands, and their fields are still unresolved, expressed by the question mark <code>?</code> in front of the name. At this point, the analyzer still has to check whether they correspond to existing index fields.</p><p>For the <code>EVAL</code> that defines <code>uppercased_mapped</code>, it can see that the previous command outputs <code>mapped_field</code>, so the unresolved <code>?mapped_field</code> marker can be replaced by a real field reference:</p><p>Next, it encounters the topmost <code>EVAL</code>, which defines <code>uppercased_unmapped</code>. The previous tree nodes produce only two fields: <code>[mapped_field, uppercased_mapped]</code>. The reference <code>?unmapped_field</code> thus has to remain unresolved. We bail here and emit the verification exception to the user.</p><h2>How unmapped_fields LOAD and NULLIFY work</h2><h3>Adding unmapped fields to the query plan</h3><p>When using <code>unmapped_fields=”NULLIFY”</code> or <code>”LOAD”</code>, we do something else; we act as if the field was actually in the index. The analyzer adds <code>unmapped_field</code> to the <code>From</code> node and marks it as unmapped to signal to the compute engine that this has to be read from <code>_source</code> or filled with <code>null</code>s. Let’s express this with a <code>{u}</code> (for <strong>u</strong>nmapped):</p><p>After amending the <code>From</code>, the analyzer can continue trying to resolve the topmost <code>Eval</code> node. It sees that the upstream nodes produce the fields <code>[mapped_field, unmapped_field, uppercased_mapped]</code> and thus <code>unmapped_field</code> can be correctly resolved:</p><p></p><p>The query plan is now fully resolved and can be passed down the regular optimization-execution pipeline. Other than the actual value extraction mechanism, everything stays the same. Schematically, the workflow looks like this:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt59bf2d2da201d00a/6a8ee8df8658b77c28469356/image2.png" alt="ES|QL unmapped fields flow: analyzer retries with NULLIFY or LOAD instead of failing on an unresolved field" /><h3>Example: enabling unmapped fields with the SET directive</h3><p>To give an example, let’s fire up a cluster and create an index with non-dynamic mappings.</p>PUT /index
{                                 
  "mappings": {
    "dynamic": false,
    "properties": {
      "mapped_field": {"type": "keyword"}
    }
  }
}

POST /index/_doc?refresh
{
  "mapped_field":"foo"
  "unmapped_field": "bar"
}<p>We can run the example query, above:</p>POST /_query
{
  "query": """
           FROM index
           | EVAL uppercased_mapped = TO_UPPER(mapped_field)
           | EVAL uppercased_unmapped = TO_UPPER(unmapped_field)
           """
}<p>This should result in the error message:</p><p><code>Unknown column [unmapped_field], did you mean [mapped_field]?</code></p><p>To make things work, we can prepend <code>SET unmapped_fields=”...”;</code> with <code>LOAD</code> or <code>NULLIFY</code>:</p>POST /_query
{
  "query": """
           SET unmapped_fields="LOAD";
           FROM index
           | EVAL uppercased_mapped = TO_UPPER(mapped_field)
           | EVAL uppercased_unmapped = TO_UPPER(unmapped_field)
           """
}

 mapped_field  |unmapped_field |uppercased_mapped|uppercased_unmapped
---------------+---------------+-----------------+-------------------
foo            |bar            |FOO              |BAR<h3>Inspecting the analyzer's rewrite steps</h3><p>If you want to see what the query analyzer is doing to the parse tree, you can log the query rewrite steps, like so:</p>PUT /_cluster/settings"
{
  "transient" : {
    "logger.org.elasticsearch.xpack.esql.analysis.Analyzer.changes": "TRACE"
  }
}<p>This will log a line containing <code>Rule rules.ResolveUnmapped applied with change…</code> You’ll see that <code>unmapped_field</code> is added to the bottom of the parse tree as described above.</p><h2>Why we have to infer the schema</h2><p>Of course, this isn’t the only possible method to deal with unmapped fields. Here are some alternatives:</p><ol><li><p>We could also scan or probe the documents in <code>index</code> to determine that their <code>_source</code> actually has the <code>unmapped_field</code>.</p></li><li><p>We could disable verifications in the analyzer and make the compute engine blindly pass unmapped fields through individual computation steps.</p></li></ol><p>The first alternative front-loads more work to understand the <em>actual</em> schema of an index and thus generally increases latency. It doesn’t scale to large, highly distributed datasets. The second alternative isn’t viable since it means a large-scale change to how ES|QL’s compute engine is built, because it passes around streams of data with fixed columns from one operator to another.</p><p>In contrast, the approach we chose is neatly compatible with ES|QL’s existing optimization pipeline.</p><p>The trade-off is that the analyzer has to correctly <em>infer</em> a schema based on the actual index mappings (obtained from the field caps endpoint) and additional fields used inside the query. </p><p>This isn’t always straightforward. There were two main challenges:</p><ol><li><p>There are many different query shapes and commands that can be used. The mechanism needs to detect unmapped fields, update the proper <code>FROM</code> command, and pass the new field through the halfway resolved plan correctly in all cases.</p></li><li><p>There are many different mappings we have to deal with, and we specifically need to make our feature work correctly when mappings <em>change over time</em> on top of that.</p></li></ol><p>In the following, we’ll focus on <code>LOAD</code>, although some problems (generally many fewer) also apply to <code>NULLIFY</code>.</p><h3>Which index to load unmapped fields from for LOOKUP JOIN and FORK</h3><p>To briefly illustrate the first problem, here are some choices we needed to make:</p><ul><li><p>Which index do we load from when using <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/lookup-join">lookup joins</a>? This one?</p><p>The <code>unmapped_field</code> cannot be attributed to both indices. We chose <code>index</code> since this is where we expect mappings to change more often than in lookup indices.</p></li><li><p>Similarly, how do we deal with subqueries and views or the <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/fork"><code>FORK</code> command</a>? In the following:</p><p>one fork branch triggers loading of an unmapped field. Is it also present in the other fork branch? (Yes, it should be, but it’s not obvious and is specifically not true if the two <code>FORK</code>s are replaced by independent subqueries.)</p></li></ul><h3>Two principles to keep queries working</h3><p>The second problem, diversity of mappings and their evolution over time, is a far bigger driver for complexity. We strove for two basic usability principles:</p><ul><li><p>Queries that work in the default mode should generally still work when using <code>unmapped_fields=”NULLIFY”</code> and <code>”LOAD”</code>.</p></li></ul><ul><li><p>Queries that work when all fields are mapped should generally still work with <code>NULLIFY</code> and <code>LOAD</code> when a field becomes unmapped and vice versa.</p></li></ul><h3>The type of unmapped fields and inadvertent type conflicts</h3><p>Let’s talk about data types to see where this leads to complexity. First, when using <code>unmapped_fields=”LOAD”</code>, we need to assume a data type for unmapped fields. We chose <code>KEYWORD</code>, which allows us to avoid type conflicts when reading from <code>_source</code>. One document can contain <code>”unmapped_field”: “foo”</code>, and another can contain <code>”unmapped_field”: 123.4</code>. It’s fine because we treat both as strings.</p><p>However, this is a violation of the second principle when a non-<code>KEYWORD</code> field happens to go unmapped. Consider this query:</p><p>If <code>some_field</code> becomes unmapped, we’ll have to assume that the <code>KEYWORD</code> type and the query will fail with a type conflict.</p><p>Type conflicts <a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-multi-index#esql-multi-index-invalid-mapping">aren’t new</a> and can be dealt with by using explicit casts in the query, like so:</p><p>It would be great if ES|QL just inferred a useful type to cast to, but this is something for the future.</p><h3>Type conflicts with partially unmapped fields, or: making PUNKs well behaved</h3><p>In addition to fully unmapped fields, <em>partially unmapped</em> fields are everywhere and should also work with <code>LOAD</code>. Let’s look at a query that uses multiple indices.</p><p>Let’s say that there are indices <code>index</code> and <code>index_without_some_field</code>, containing just the following documents.</p>// index1
{
  "some_field": "foo"
}

// index2
{
  "some_field": "bar"
}<p>Now let’s consider the query:</p>FROM index, index_without_some_field<p>and assume that <code>some_field</code> is unmapped in <code>index_without_some_field</code>. This will return:</p>some_field
-------------
 foo
 null<p>because ES|QL doesn’t load unmapped fields per default.</p><p>Of course, when setting <code>unmapped_fields=”LOAD”</code>, we want to load from <code>_source</code> for <code>index_without_some_field</code>:</p>SET unmapped_fields="LOAD";
FROM index, index_without_some_field

 some_field
-------------
 foo
 bar           // loaded from _source<p>As with fully unmapped fields, the case is simple when <code>some_field</code> is mapped as <code>KEYWORD</code> in <code>index</code>. When loading from <code>_source</code> for <code>index_without_some_field</code>,  we treat the field as <code>KEYWORD</code> as well, so there’s no conflict.</p><h3>What makes a field a PUNK</h3><p>The case is less clear when <code>some_field</code>is partially unmapped and the mapped leg is of a type other than <code>KEYWORD</code>. Such fields caused a lot of trouble until we found the best solution, which makes their acronym quite fitting: <strong>p</strong>artially <strong>u</strong>nmapped <strong>n</strong>on-<strong>k</strong>eyword fields, or PUNKs.</p><p>Unfortunately, PUNKs are far from being esoteric. For instance, it’s very natural to filter on a PUNK:</p><p>If <code>some_field</code> is mapped as <code>INTEGER</code> in <code>index</code>, the type conflict looks like this:</p><ul><li><p>Mapped as an <code>INTEGER</code> in <code>index</code>.</p></li><li><p>Unmapped in <code>index_without_some_field</code> and thus treated as <code>KEYWORD</code>.</p></li></ul><p>This can again be resolved manually by providing an explicit cast:</p><p>But this is far from acceptable. Even queries that work fine without <code>NULLIFY</code> and <code>LOAD</code> typically have <em>some</em> PUNKs; the unmapped leg is simply treated as <code>null</code> then. Both guiding principles are violated if <code>LOAD</code> requires an explicit cast here.</p><h3>Casting implicitly to the mapped type</h3><p>The solution is to introduce an implicit cast to the mapped type. In this case, we know that <code>some_field</code> is an <code>INTEGER</code> in <code>index</code>, and thus we treat it essentially as if the user wrote:</p><p>This means that queries that work without <code>LOAD</code> keep working. (ES|QL may even give you more data because we load the unmapped leg of PUNKs from <code>_source</code>.) Queries that used to work when a field is fully mapped also keep working when it goes unmapped in some (but not all) of its indices without having to alter the query in any way.</p><p><strong>Behavior</strong></p><p><strong>Default</strong></p><p><strong><code>NULLIFY</code></strong></p><p><strong><code>LOAD</code></strong></p><p>Unmapped field in query</p><p>Query fails</p><p>Query runs</p><p>Query runs</p><p>Values returned</p><p>None</p><p><code>null</code></p><p>Read from <code>_source</code> </p><p>Assumed type</p><p>n/a</p><p><code>NULL</code></p><p><code>KEYWORD</code></p><p>Partially unmapped field (PUNK)</p><p>Unmapped leg is <code>null</code></p><p>Unmapped leg is <code>null</code></p><p>Cast to the mapped type</p><p>Pushdown optimization</p><p>Full</p><p>Full</p><p>Per-node where fully mapped</p><h2>Don't throw it all away: Keeping ES|QL query optimization with unmapped fields</h2><p>There's one more thing to get right; that is, to make sure that optimizations still work correctly with <code>LOAD</code>. Consider the previous query:</p><p>ES|QL’s optimizer aggressively pushes down such <code>WHERE</code> filters and turns them into Lucene queries, so the compute engine doesn’t perform unnecessary work.</p><p>For this query, evaluating the filter in the compute engine would require fetching each and every document from the index; meaning, a full scan, very slow. If <code>some_field</code> was mapped as an <code>INTEGER</code> in both indices, we would instead perform a Lucene query, which looks like this:</p>{
  "range": {
    "some_field": {
      "gt" : 10,
      "boost" : 0.0
    }
  }
}<p>The compute engine then doesn’t have to load each document separately and check whether it matches the filter. Documents with <code>some_field &lt;= 10</code> are never fetched from the Lucene index, which is very efficient at this kind of filtering. Nice.</p><h3>Why filter pushdown is unsafe for unmapped fields</h3><p>If <code>some_field</code> is unmapped in <code>index_without_some_field</code>, however, it’s wrong to narrow the documents down using the same Lucene query, as Lucene interprets an unmapped <code>some_field</code> as <code>null</code> and thus no documents from <code>index_without_some_field</code> will ever match. This edge case is easy to miss, and it doesn’t help that there are several flavors of similar pushdowns. For instance, in the query:</p><p>the compute engine pushes even the counting to Lucene. Again, this is only correct if <code>some_field</code> is fully mapped.</p><p>This means that such optimizations can't apply to unmapped fields. It would be disappointing if a query used hundreds of indices and only one of them happened to not map <code>some_field</code>, causing the whole query to run unoptimized.</p><h3>How the local optimizer recovers the fast path</h3><p>Luckily, this problem has a solution, too. ES|QL actually has multiple optimizer runs:</p><ol><li><p>First, a preliminary optimizer run on the node handling the <code>_query</code> request.</p></li><li><p>Then, a second, local optimizer run on every node we fan out to because we need to fetch documents from its shards.</p></li></ol><p>The workflow after the initial optimization looks more like this:</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb2c907ea31960f96/6a8eeabe36492afa7685fa36/image1.png" alt="" /><p>If the current node happens to map <code>some_field</code> in all shards, the local optimizer detects this situation and treats <code>some_field</code> like any other fully mapped field, including performing Lucene queries to greatly narrow down the dataset to be processed. In fact, data nodes process <code>LIMIT</code> queries like:</p><p>in batches of shards (to avoid loading too much data too eagerly), which includes a full local optimizer run per batch. This makes it even more likely to encounter batches where <code>some_field</code> is fully mapped, allowing ES|QL to run a fast Lucene query.</p><h2>Is it working now? Testing unmapped_fields across every ES|QL query shape</h2><p>As we have seen from the optimizer issues above, problems can hide in plain sight, even for very simple queries. Because <code>unmapped_fields=”LOAD”</code> can affect each and every kind of query, the surface area for bugs is essentially all of ES|QL.</p><p>Accordingly, getting good test coverage was tricky and challenged us to refine our testing strategies.</p><h3>Reusing spec tests with unmapped_fields</h3><p>Conveniently, ES|QL has an extensive corpus of test queries, together with expected result sets; we call them <em>spec tests</em> because they’re written using a simple text specification language, which looks roughly like this:</p>simpleEval
row a = 1 | eval b = 2
;

a:integer | b:integer
1         | 2
;<p>This lets us create new tests out of the existing ones by introducing slight variations. For instance, any existing test that runs without <code>SET unmapped_fields=”...”</code> should produce the exact same results when run with <code>SET unmapped_fields=”NULLIFY”</code>.</p><p>It also helped find major issues early in the development process, especially for <code>NULLIFY</code>. The <code>LOAD</code> setting changes the meaning of queries much more dramatically, limiting the usefulness of this approach. However, ES|QL also uses what we call <em>generative testing</em>; that is, we string together random commands, run the query, and then check whether the server reports a bug. This approach cannot confirm the correctness of results, but it still helped greatly with finding query types that didn’t work properly and resulted in some kind of error. (Property-based tests would be a refinement in the future by running the queries against a reference implementation. This way, correctness of results can also be checked.)</p><h3>Testing type conflicts across different mappings</h3><p>In the end, one of the most important testing dimensions was using different indices with various mappings in the same query. (Recall how, above, we had to deal with type conflicts to come up with a solid approach for PUNKs? It doesn’t end there; all kinds of type conflicts are more complex with <code>LOAD</code>.) Since we couldn’t automatically generate correct expected results, ES|QL’s test suite had to grow by adding more than 10,000 lines of CSV spec tests. Fortunately, adding such tests is a well-suited task for an AI agent, which has cut down the effort dramatically. (Of course, the test results were still reviewed by humans.)</p><p>All testing strategies together provided us with good confidence for the GA release of <code>unmapped_fields</code> with Elasticsearch 9.5.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/esql-unmapped-fields-deep-dive</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/esql-unmapped-fields-deep-dive</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Mappings]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <dc:creator><![CDATA[Alexander Spies]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9f8079c053533ad8/6a8ee79f8658b748a0469342/image4.png" length="0" type="image/png"/>
    <pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Skip the mapping explosion: ES|QL queries schemaless JSON keys without dynamic mapping]]></title>
    <description><![CDATA[Flattened fields turn Elasticsearch into a schema-on-read store where you index schemaless data under one mapping, then use ES|QL's FIELD_EXTRACT to pull out any JSON key you need to filter, group or join on, with predicates pushed into the columnar store.]]></description>
    <content:encoded><![CDATA[<p><a href="https://www.elastic.co/docs/reference/query-languages/esql">ES|QL</a> now reads<a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/flattened"> <code>flattened</code> fields</a>. FIELD_EXTRACT pulls any key out of a schemaless JSON object so you can filter, group, sort and join on keys you never mapped. The planner pushes those predicates into the columnar store rather than parsing the whole blob per row, which means dynamic JSON keys from OTel attributes, log labels, user metadata or whatever else you didn't want to map individually are queryable without causing a<a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/mapping-limit"> mapping explosion</a>.</p><h2>Why dynamic mapping breaks down with schemaless data</h2><p>Elasticsearch wants a schema. Every field you index has a mapping that defines its type, its analyzer (for text), and how it is stored. That schema makes storage compression efficient, with fast search and cheap aggregation. It becomes a liability the moment your data stops looking like a database table.</p><p>Consider the shapes that show up in real systems:</p><ul><li><p>Log events where every service adds its own attributes. One emits <code>labels.region</code>, another <code>labels.k8s.pod</code>, and another <code>labels.tenant_id</code>.</p></li><li><p>OpenTelemetry resource attributes, where the set of keys is defined by whatever agent happened to send the span.</p></li><li><p>User-supplied metadata bags, feature flags, or tagging systems, where the keys are open-ended by design.</p></li></ul><p>If you map each key as its own field, the mapping grows unbounded. This is the classic "mapping explosion." Thousands of dynamically created fields inflate the cluster state, slowing down every mapping update and eventually hitting the field limit. Each field also carries overhead in the index. You pay a structural cost for keys you didn’t plan for and may query only once.</p><p>The naive escape hatch is to store the whole object as a string and give up on querying its contents. That trades one problem for another: you keep the data but lose the ability to filter or group by anything inside it.</p><p>The <code>flattened</code> field type is the better path. You can index an entire JSON object under a single mapped field, keeping the keys queryable with almost none of the mapping-explosion cost. </p><p>This post covers how a flattened field stores the data on a disk and how ES|QL reads it back.</p><h2>How flattened fields index dynamic JSON keys under one mapping</h2><p>Map one field as flattened:</p>PUT logs
{
 "mappings": {
"properties": {
"labels": { "type": "flattened" }
   }
 }
}<p>Then write arbitrary nested JSON into it:</p>POST logs/_doc
{
 "labels": {
   "region": "us-east-1",
   "k8s": { "pod": "web-7f9", "node": "ip-10-0-0-3" },
   "retries": 4
 }
}<p>There is exactly one field in the mapping, <code>labels</code>, no matter how many keys appear across your documents. The cluster state doesn’t grow when a new key shows up. The subkeys remain individually searchable. You can reference <code>labels.region</code> or <code>labels.k8s.pod</code> in queries, even though neither was ever declared.</p><p>The catch and central tradeoff is that every leaf value is a keyword. The number 4 above is indexed as the string <code>"4"</code>. There is no numeric typing, no date parsing, and no range math on dynamic keys. Flattened fields exchange per-field richness for schema flexibility. That fact explains almost every design decision that follows.</p><h2>How Elasticsearch stores flattened field JSON keys on disk</h2><p>There are two types of queries on flattened fields: an unkeyed query on the root flattened field, and a keyed query on a specific subfield. Following the previous example, a query of the form <code>labels: "us-east-1"</code> matches a value under <em>any</em> key, while the query<code>labels.region: "us-east-1"</code> matches a value under the <em>specific</em> <code>region</code>key.</p><p>To support these two distinct query formats, the flattened mapper writes each leaf value into two distinct Lucene fields.</p><p>Take this document:</p>{ "labels": { "region": "us-east-1", "k8s": { "pod": "web-7f9" } } }<p>The mapper produces:</p><ul><li><p>A root field under labels, holding the bare values:</p></li></ul>us-east-1
web-7f9<ul><li><p>A keyed field under labels._keyed, holding the flattened key concatenated with its value:</p></li></ul>region\0us-east-1
k8s.pod\0web-7f9<p>In the keyed field, nested objects are dot-flattened into a single key (k8s.pod), and the key is joined to its value with a reserved NULL byte (\0) as the separator. Keys that contain a NULL byte are rejected at parse time, so the separator is always unambiguous. To find the value, you split on the first NULL.</p><p>These two fields make both query shapes work:</p><ul><li><p><code>labels: "us-east-1"</code> matches a value under <em>any</em> key, so it searches the root field.</p></li><li><p><code>labels.region: "us-east-1"</code> matches a value under a <em>specific</em> key. It rewrites the query to the term <code>region\0us-east-1</code> and searches the keyed field.</p></li></ul><h2>Query restrictions on flattened field subkeys</h2><p>Because every key's terms live in one sorted list, the keyed field cannot answer every query shape a plain keyword field can. Three restrictions follow:</p><ol><li><p>No fuzzy, regexp, or wildcard on a specific subkey. Nothing about the layout makes them impossible,  but the pattern would have to be combined with the key prefix so it can’t walk past the key boundary. The mapper doesn’t do that at this writing.</p></li><li><p>Any range query on a subkey needs at least one bound. Elasticsearch already has a query for "this field has some value here, whatever it is": the <code>exists</code> query. On a flattened subkey it runs as a prefix query on key\0, which sweeps every term belonging to that key. A range with neither bound would sweep exactly the same terms. Rather than support two spellings of one scan, the mapper rejects the boundless range and asks for the <code>exists</code> query.</p></li><li><p>A range query on a subkey needs the field to be indexed. A flattened field can be mapped with <code>index: false</code>, which skips the inverted index and keeps only doc values, the columnar per-document storage covered in the next section. Exact-match queries survive that. With no terms to look up, Elasticsearch scans the doc values column instead, which is slower but gives the same answer. Range queries have no equivalent fallback, so a range on a subkey of an unindexed flattened field throws an Exception.</p></li></ol><p>All three are limits on the Lucene query the mapper is willing to build, and where you notice them depends on how you query.</p><p>On the search API, where you name <code>labels.region</code> directly, they come back as errors.</p><p>In ES|QL you will not see them as errors at all. There, the same limits decide only whether a predicate is pushed into Lucene or runs in the compute engine on the extracted column. </p><p>This is pushed to a term query on the keyed field:</p><p>This is not, so the filter runs per row on the extracted keyword:</p><p>Same answer, more work. That distinction is the subject of the second half of this post.</p><h2>Under the hood: how range queries stay inside key boundaries</h2><p>This part is internal. You don’t need it to use the field, but it explains where the bounds rule comes from.</p><p>For a handful of documents in one segment, the shared term list looks like this:</p>k8s.pod\0web-7f9
region\0us-east-1
region\0us-west-2
tenant_id\0acme<p>Each key owns a contiguous slice of that list. For example, all values for key "region" are clustered together in an ordered sublist. A range with both bounds set encodes each bound the same way a term is encoded, so a lower bound of "us-east" on the region key becomes region\0us-east and an upper bound of "us-west" becomes region\0us-west. Both endpoints already carry the key prefix, so the scan can’t leave the region slice. Nothing special is needed.</p><p>The half-open case is problematic. Handing Lucene a lower bound of <code>region\0us-east</code> with no upper bound would scan to the end of the term list, straight through <code>tenant_id\0acme</code> and every other key that sorts after region. So the mapper substitutes a sentinel for the missing side:</p><ul><li><p>A missing lower bound becomes key\0, inclusive. That’s the encoding of the empty value, and it’s the first term in the key's slice.</p></li><li><p>A missing upper bound becomes key\1, exclusive. Byte 0x01 is the next byte after the 0x00 separator, so it sorts after every key\0value term and before the first term of any other key.</p></li></ul><p>A one-sided range is therefore boxed into [key\0, key\1), which makes it exactly as safe as a closed one.</p><p>The upper sentinel also covers the case where one key is a prefix of another. If an index holds both region and regionx, then region\1 still sorts below regionx\0eu-west-1, because 0x01 is smaller than the x that follows the shared region prefix. A range on region cannot leak into regionx.</p><h3>How flattened fields use the inverted index and doc values</h3><p>Each leaf value can be written into two Lucene structures. Both are enabled by default, but can be disabled by the mapping configuration.</p><ul><li><p>The inverted index (when the field is indexed). 
Two untokenized keyword terms are indexed per value, one on the root path and one on the keyed path. This powers term, prefix, and range searches.</p></li><li><p>Doc values (when <code>doc_values</code> is enabled). 
A columnar, document-ordered structure. This powers sorting, aggregations, and ES|QL reads.</p></li></ul><p>The inverted index answers the question: "Which documents contain this term?" </p><p>Doc values answer another question: "For this document, what are the values?" Doc values are laid out column by column so a scan touches only the bytes it needs. A flattened field uses both inverted indexes and doc values, so it can serve search and analytics from the same field.</p><p>The nature of the index means that the root field is only present in some cases. When the inverted index is disabled by the mapping, any value search requires a linear scan of the doc values for the searched value. In this case, when performing a search on the root field, there isn’t much additional overhead compared to just scanning the keyed field and ignoring the key markers. So the flattened mapper skips writing the root field, relying on the keyed field for both root and keyed queries.</p><h3>Why flattened fields switched from dictionary to binary doc values</h3><p>Historically, flattened field doc values used Lucene's <code>SortedSetDocValues</code>, which is a dictionary-compressed format. This means that every <code>key\0value</code> value indexed across all documents per segment is stored in one big sorted, deduplicated set of values. Each document tracks a list of ordinals into that value set.</p><p>This dictionary approach provides great compression for low-cardinality fields that tend to have repeated values. It is byte-efficient to store a value only once, and then refer to it by a single integer value. However, there is overhead associated with building and maintaining that dictionary of values, and that overhead is wasted effort when operating on high-cardinality fields that don’t repeat values.</p><p>Because flattened fields are a catch-all type, their cardinality tends to be very high. So while the dictionary approach works, it’s not the most compact or scan-friendly layout for this data, especially in time-series indices where flattened bags are common and storage pressure is real.</p><p>Recent versions switched the storage to use Lucene’s <code>BinaryDocValues</code>. This format just stores a literal binary blob for each document, which is compressed using Zstandard by our doc values codec when written to disk.</p><p>This new binary format provides an additional benefit: it allows us to maintain original array ordering without any overhead. Flattened fields support the mapping parameter <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/flattened#flattened-params">preserve_leaf_arrays</a>, which affects how multivalued fields are returned when using <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/mapping-source-field#synthetic-source">synthetic source</a>. When configured to <code>preserve_leaf_arrays: exact</code>, returned values preserve the order, duplicates, and nulls from the original source value.</p><p>The dictionary encoding inherent to sorted-set doc values means the returned values are sorted, deduplicated, and de-nulled. To implement <code>preserve_leaf_arrays</code>, flattened fields have traditionally used an additional sidecar field, tracking the required metadata to reconstruct the original source value. However, the nature of binary doc values means that this sidecar field is no longer needed. The values are just stored and returned as indexed.</p><h3>Limitations of schema on read with flattened fields</h3><p>The limitations below follow from the keyword-only rule and the shared keyed field:</p><ul><li><p>No numeric, date, or Boolean typing on dynamic keys. <code>100</code> and <code>"100"</code> are the same term.</p></li><li><p>No fuzzy, regexp, or wildcard on a specific subkey.</p></li><li><p>No multi-fields (<code>fields</code>) or <code>copy_to</code> on the flattened field.</p></li><li><p>A <code>depth_limit</code> (default 20) on how deeply nested the object can be.</p></li></ul><p>If you need real typing for a <em>known</em> key, flattened now supports explicitly <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/flattened#flattened-properties">mapped subfields</a>: you declare individual keys with real types under a <code>properties</code> block, and those keys are indexed by their own typed mapper instead of the keyed field. You get numeric ranges on <code>labels.status_code</code>, while everything else in <code>labels</code> stays dynamic and keyword-only. This is the escape hatch for the handful of keys you actually know about.</p><h2>Schema on read: querying flattened field JSON keys in ES|QL</h2><p>Search has supported flattened fields for years. ES|QL, the newer, piped query language built on a columnar compute engine, now reads flattened fields, too. Support began in the Technical Preview of Elasticsearch 9.5.0. It has two distinct pieces.</p><h3>What ES|QL returns when you select a flattened field</h3><p>ES|QL uses a dedicated data type for flattened fields, rather than folding it into <code>keyword</code>. When you select the root, you get the whole object back as a JSON string:</p>labels:flattened
{"k8s.pod":"web-7f9","region":"us-east-1"}<p>Keys come back sorted. You can carry this value through a query, count it, group by it, and run the multi-value and comparison functions on it. But an opaque JSON blob is not usually what you want to filter or aggregate on. For that, you need to reach inside it and process its contents.</p><h3>How FIELD_EXTRACT reads JSON keys from flattened fields</h3><p>There is no dotted-path syntax for dynamic keys in ES|QL. You cannot write <code>labels.region</code> for an unmapped key, because to the engine the flattened root is a single leaf value, not a set of columns. Instead you use a function:</p><p>The absence of a dotted-path syntax is a tentative limitation. The keyed field already addresses individual subkeys, so a more natural syntax for reaching into a flattened root is something we are planning to support.</p><p>FIELD_EXTRACT(field, path) takes a flattened field and a key, and returns a keyword. The rules are easier to see against a document. Take this one:</p>POST logs/_doc
{
 "labels": {
   "region": "us-east-1",
   "k8s": { "pod": "web-7f9", "node": "ip-10-0-0-3" },
   "tags": ["prod", "canary"],
   "retries": 4
 }
}<p>And this query:</p><p>The result is:</p>region     | pod     | k8s  | tags            | retries | namespace
us-east-1  | web-7f9 | null | [prod, canary]  | 4       | null<p>Four things to take from this:</p><ul><li><p>The dot is part of the key, not a navigation operator. The mapper already collapsed the nested object into the flat key k8s.pod, so "k8s.pod is a direct lookup, not a walk from k8s to pod.</p></li><li><p>Matching is exact. "k8s" returns null because there is no leaf stored at k8s, only at k8s.pod and k8s.node. For the same reason, "host" will not find "host.name", and matching is case-sensitive, so "Region" will not find "region".</p></li><li><p>Arrays come back multi-valued. "tags" yields a multi-valued keyword you can <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/mv_expand">MV_EXPAND</a>, count, or filter on. A missing key yields null.</p></li><li><p>Everything is a keyword. "retries" comes back as the string "4", and a Boolean leaf comes back as "true" or "false".</p></li></ul><p>JSONPath syntax is rejected outright, at parse time rather than per row. Both FIELD_EXTRACT(labels, "['k8s.pod']") and FIELD_EXTRACT(labels, "tags[0]") fail with <em>field_extract path must be a literal flattened sub-field name</em>.</p><p>Once extracted, the value is an ordinary <code>keyword</code> column. You can filter on it, group by it, sort by it, or use it as the join key in a <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/lookup-join"><code>LOOKUP JOIN</code></a>:</p><h3>How ES|QL pushes flattened field predicates into the columnar store</h3><p>The obvious way to implement FIELD_EXTRACT would be to read the whole flattened root out of storage, render it as a JSON string, hand it to the compute engine, and parse it once per row to pull out a single key. But that means reading every key to use once and paying for a JSON parse on every document. So this obvious implementation is not performant.</p><p>ES|QL avoids this whenever possible. Between storage and the compute engine sits the <em>block loader</em>, the step that turns stored data into the columnar blocks the engine operates on. FIELD_EXTRACT hooks into that step instead of running after it.</p><p>When ES|QL loads a column, the flattened field type inspects the request. If the request is an extraction of a single constant key and the field has doc values, it routes straight to the keyed doc-values loader, which reads the key\0value entries for only that key out of the columnar structure. The column that arrives at the compute engine already contains only that key's values. The JSON string is never built and never parsed, and the other keys in the object are never read.</p><p>When extraction can’t be fused into the block loader, for example, because the key is computed per row or the root is the output of another function such as CASE, ES|QL falls back to the parse-per-row path. The results are identical either way. Only the cost changes.</p><p>The comparison can push down further. A predicate like <code>FIELD_EXTRACT(labels, "region") == "us-east-1"</code> can be pushed to Lucene as a term query against the synthetic keyed field, the same <code>region\0us-east-1</code> term the search path uses. So a filter on an extracted subkey can be answered by the inverted index (if available), and the projection can be answered by doc values, exactly like a first-class field, even though the key was never in the mapping.</p><p>Ordering comparisons push down, too. The four, single-sided comparators (&gt;, &gt;=, &lt;, &lt;=) and closed BETWEEN-style ranges all become a range query on that same synthetic keyed field. This is where the key\0 / key\1 sentinels earn their keep: the single-sided forms are only pushable because the mapper can box an open bound inside the key. The pushed range is treated as a candidate, and the predicate is re-evaluated on the extracted keyword column afterwards, so multi-valued keys don’t slip through.</p><p>The values are keywords, so the ordering is lexicographic, not numeric. FIELD_EXTRACT(labels, "retries") &gt; "10" compares strings, which means "9" is greater than "10". If you need numeric ranges on a key, map it explicitly under properties, or cast the value in ESQL.</p><p>Explicitly mapped subfields behave differently on purpose. Because they carry real types, comparison semantics diverge from the keyword path, so they are loaded and compared through their own typed mapper rather than fused into the keyed loader. And when you select a flattened root that has mapped subfields, ES|QL loads it from <code>_source</code>, so every leaf renders as a string and no keys are dropped silently.</p><h2>When to use flattened fields vs. dynamic mapping in Elasticsearch</h2><p>Use <code>flattened</code> fields when:</p><ul><li><p>The set of keys is open-ended or unknown ahead of time.</p></li><li><p>You would otherwise cause a mapping explosion.</p></li><li><p>Keyword-level filtering and grouping on the values is enough, and you don’t need numeric or date semantics on the dynamic keys.</p></li><li><p>You have a few keys that <em>do</em> need real types. Map those explicitly under <code>properties</code>, and let the rest stay dynamic.</p></li></ul><p>Avoid it, or map fields normally, when the schema is stable and you need full-text analysis, numeric aggregation, or date math across the board.</p><p>Remember that flattened is not a dumping ground for JSON you have given up on. It’s a real columnar-and-inverted store for schemaless data, and with ES|QL support, it’s now a first-class analytical citizen. You can keep the messy, unmapped parts of your data messy, and still filter, group, join, and aggregate across them as needed.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/schema-on-read-esql-json-keys</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/schema-on-read-esql-json-keys</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Mappings]]></category>
    <category><![CDATA[Lucene]]></category>
    <dc:creator><![CDATA[Jordan Powers,Dima Leontyev]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltb3fca0aedcfba9ab/6a8419fae41d7f68b46522b9/unnamed.png" length="0" type="image/png"/>
    <pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Elasticsearch ES|QL brings full-text search to data you never indexed]]></title>
    <description><![CDATA[MATCH and TO_TEXT bring full-text search to data you never indexed. Search computed columns, unmapped fields and federated sources in ES|QL.]]></description>
    <content:encoded><![CDATA[<p>ES|QL <code>MATCH</code> now runs full-text search on data you never indexed. Computed columns, unmapped fields, strings assembled on the fly, even federated data sitting in S3. The new <code>TO_TEXT</code> function tells ES|QL to treat any string as analyzable text, so <code>MATCH</code> can tokenize, case-fold and term-match values that exist only for the lifetime of a query. This goes beyond the <code>LIKE</code> and <code>RLIKE</code> pattern matching that most query engines offer for unindexed strings: it's real analysis. Available now in Elastic Cloud Serverless and as a technical preview in Elasticsearch 9.5.</p><h2>How MATCH and TO_TEXT enable full-text search on any ES|QL expression</h2><p>Let's start with a query that was impossible in Elasticsearch 9.4, which uses <a href="https://www.elastic.co/docs/reference/query-languages/esql/commands/eval">the <code>EVAL</code> command</a>:</p><p>In this example, <code>summary</code> has no mapping or analyzer configuration. It’s also not associated with any inverted index. It exists only for the lifetime of this query, but now you can search it anyway. Two additions make this work.</p><p>First, <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/search-functions/match"><code>MATCH</code></a> now accepts any expression as its first argument, not just a mapped field. That includes columns produced by <code>EVAL</code> and function results used inline. It also includes unmapped fields loaded directly from the original document. Furthermore, all data types normally accepted by <code>MATCH</code> are supported in this new use case.</p><p>The second part of this is the new <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/type-conversion-functions/to_text"><code>TO_TEXT</code></a> function, which is the first ES|QL conversion function that produces output of type <code>text</code>. Until now, <code>text</code> columns could only come from indexed mapped fields, and all strings produced by ES|QL expressions were <code>keyword</code> values rather than <code>text</code>. The distinction matters because <code>MATCH</code> treats the two differently: <code>text</code> values are analyzed, while <code>keyword</code> values are compared exactly, mirroring how a <code>MATCH</code> query on an indexed keyword field rewrites to a term query. <code>TO_TEXT(x)</code> is how you tell ES|QL: <em>treat this string as full text</em>.</p><p>This ships as a technical preview in Elasticsearch 9.5, and as such, it has some limitations:</p><ul><li><p>It’s currently filtering only. A <code>MATCH</code> on an expression doesn't contribute to the relevance score yet; only matches on indexed fields affect the score.</p></li><li><p>Querying options like <code>fuzziness</code> and others aren't yet supported when matching an expression.</p></li><li><p>Runtime text is analyzed with the standard analyzer. This isn’t configurable yet.</p></li></ul><p>Work is underway to address these limitations.</p><h2>Why use full-text search instead of LIKE or RLIKE in ES|QL?</h2><p>ES|QL already had two ways to search strings without an index: <code>LIKE</code> (wildcard patterns) and <code>RLIKE</code> (regular expressions). Both work on any string expression, so it's fair to ask what <code>MATCH</code> adds. The answer is <a href="https://www.elastic.co/docs/manage-data/data-store/text-analysis">analysis</a>, a more advanced form of search which uses techniques such as stemming and synonyms. It also uses stopword handling.</p><p><code>LIKE</code> is simple substring matching, without any understanding of the words that comprise a string. Say, for example, you're looking for log messages about a fox:</p><p>This misses <code>"Fox spotted near the henhouse"</code> due to the capitalization, while matching <code>"Outfoxed by the competition"</code>, which isn't about a fox at all. It fails in both directions, with false negatives on capitalization and false positives on substrings buried inside other words.</p><p>Regular expressions can patch the case problem, but the word-boundary problem gets ugly fast. Something like:</p><p>And even that's not right yet. It misses a fox at the end of a sentence followed by <code>!</code> or <code>?</code>, and it says nothing about tabs, quotes, or parentheses. Each fix makes the pattern longer, and the next person to read the query has to reverse-engineer what it’s actually doing.</p><p><code>MATCH</code> makes the problem go away, because it runs both the query and the value through an <a href="https://www.elastic.co/docs/reference/text-analysis/analyzer-reference">analyzer</a>, which tokenizes text into lowercase terms and then matches term against term:</p><p>This query will match values like <code>"The quick brown fox"</code> and <code>"FOX spotted near the henhouse"</code> but not <code>"Outfoxed by the competition"</code> or <code>“FOXTROT protocol enabled"</code>, regardless of any punctuation surrounding the words. Of course, this all works for multi-term queries, like <code>MATCH(TO_TEXT(message), "brown fox")</code>, too, just the way you’d expect it to.</p><p>Work is underway to enable the use of the <a href="https://www.elastic.co/docs/reference/text-analysis/analysis-lang-analyzer">36 dedicated language analyzers</a>, with support for natural languages on data that was never indexed or mapped.</p><h2>Full-text search use cases for unindexed and unmapped data</h2><p>The examples above searched values computed from <a href="https://www.elastic.co/docs/manage-data/data-store/mapping">mapped fields</a>. . The more interesting use cases for ES|QL <code>MATCH</code> on expressions involve data that was never searchable at all. Let's walk through a few.</p><h3>How to search unmapped fields in ES|QL without adding a mapping</h3><p>Sometimes you deliberately leave a field out of your mappings, such as a verbose stack trace or a raw request payload. You might even leave out a debug blob. Indexing one of these would cost disk and heap space on every document and wouldn’t be worth it for a field you might query once a quarter.</p><p>That decision has always been final, because <a href="https://www.elastic.co/search-labs/blog/esql-unmapped-fields">unmapped fields</a> were invisible to queries entirely. In Elasticsearch 9.5, you can use <code>SET unmapped_fields="load"</code> to make ES|QL load unmapped fields directly from the source document as keywords. Follow that up by wrapping it in <code>TO_TEXT</code>, and now you can run full-text search on it:</p><p>Here, <code>stack_trace</code> was never mapped. Every value is fetched from the original documents and analyzed on the fly. They’re matched row by row. That’s real work, and it will never be as fast as an <a href="https://www.elastic.co/docs/manage-data/data-store/index-basics">inverted index</a> lookup. But now, that field you didn't index is no longer unsearchable. You get to keep the mapping small for the everyday case and still answer the once-a-quarter question when it matters.</p><h3>Full-text search on a keyword field without reindexing</h3><p>Keyword fields can do a lot. They give you exact matching, fast aggregations, and sorting, which is why so many fields end up mapped that way. But mappings are decided when data arrives, and it’s easy to end up in a situation where you want to do something different with your data than you had originally intended. Maybe <code>product_name</code> was mapped as a <code>keyword</code> because the dashboards aggregate on it, and then, after receiving a year’s worth of product data, someone wants to be able to search within <code>product_name</code> values.</p><p>The old answer was to change the mapping to <code>text</code> (or add a multi-field) and reindex everything. This can be both time-consuming and costly, and in many cases, users simply won’t want to bother with it. The new answer is one function call:</p><p><code>TO_TEXT</code> converts the <code>keyword</code> values to <code>text</code> on the fly, so <code>MATCH</code> analyzes them instead of comparing them exactly. This allows you to query a <code>keyword</code> field without creating a mapping or reindexing the source document. If the search becomes an everyday query, indexing the field as <code>text</code> is still the right long-term move, but <code>TO_TEXT</code> gets you an answer today, without any extra work.</p><h3>Searching the same field across indices with different mappings</h3><p><a href="https://www.elastic.co/docs/reference/query-languages/esql/esql-multi-index">ES|QL can span many indices</a>, and the same field doesn't need to always look the same in all of them. When the same field has different types in different indices, ES|QL treats it as a union type, and a conversion function resolves the conflict. Let’s consider an example in which the <code>message</code> field has type <code>text</code> in this year's index template but was a <code>keyword</code> in last year's:</p><p>Every value is analyzed at query time, whether it came from the <code>text</code> index or the <code>keyword</code> one. The keyword values from the older indices are tokenized and lowercased like everything else, so "connection reset" finds “Connection RESET by peer”, no matter which index it lives in.</p><p>Another interesting case is when a field is mapped in only one index but also present (and unmapped) in the other:</p><p>There's a nuance worth calling out here. If <code>error_details</code> is mapped in <code>logs-2026</code> but not in <code>logs-2025</code>, Elasticsearch cannot push this query down to <a href="https://lucene.apache.org/">Lucene</a>, because the indices where the field is unmapped would silently return no matches. Instead, the planner notices that the field is potentially unmapped and evaluates the whole <code>MATCH</code> row by row, wherever the rows came from. You don't have to know which of your indices have the field mapped; the query just answers the question.</p><h2>How ES|QL analyzes text at query time without an inverted index</h2><p>When ES|QL plans a <code>MATCH</code> against an expression, it analyzes the query string once, up front, into a set of terms. How each row is then evaluated depends on the expression's type:</p><p><strong>Expression type</strong></p><p><strong>Processing</strong></p><p><strong>Matching behavior</strong></p><p><code>text</code> (via <code>TO_TEXT</code>)</p><p>Analyzer tokenizes value into lowercase terms</p><p>Token-against-token comparison; a row matches if any token equals any query term (<code>OR</code> semantics)</p><p><code>keyword</code>,<code>ip</code>, <code>date</code>, numeric</p><p>No analysis; query constant converted once to the native type</p><p>Exact comparison per row</p><p>Both paths bypass Lucene entirely and evaluate values row by row. The non-text path mirrors exactly what a match query does when pushed down to Lucene against those field types, so the semantics stay consistent regardless of whether your query hits an index.</p><p>An inverted-index lookup does its work at ingest time and never touches non-matching documents at query time. A runtime <code>MATCH</code> does that analysis at query time, for every row that reaches it. One is fast because the work already happened; the other is flexible because the data doesn't need to have been indexed at all.</p><h2>What's next for ES|QL full-text search</h2><p>Everything in this post is the first installment of a larger effort to make search in ES|QL work on anything, not just on what you indexed ahead of time. The limitations called out earlier are actively being worked on, and the roadmap goes further:</p><ul><li><p><strong>Scoring.</strong> Runtime matches will contribute to <code>_score</code>, so you can sort by relevance even when the data was never indexed.</p></li><li><p><strong><code>MATCH_PHRASE</code></strong><strong> on expressions.</strong> Already available in Elastic Cloud Serverless, and coming to the Elastic Stack in 9.6.</p></li><li><p><strong>Configurable analyzers.</strong> Analyzer support for <code>MATCH</code> and <code>MATCH_PHRASE</code> on expressions, enabling language analyzers, stemming, and synonyms at query time.</p></li><li><p><strong>Match options.</strong> Options like <code>fuzziness</code> and <code>operator</code> for runtime matches.</p></li><li><p><strong>Vector search.</strong> Generating embeddings per row and running k-nearest neighbors (kNN) on runtime <code>dense_vector</code> expressions, bringing semantic search to unindexed data, too.</p></li></ul><h2>Try ES|QL full-text search on expressions today</h2><p>You can try runtime search today. It's available now in Elastic Cloud Serverless, where new ES|QL capabilities land first, and it ships as a technical preview in Elasticsearch 9.5. Start with the <a href="https://www.elastic.co/docs/reference/query-languages/esql/functions-operators/search-functions">search functions</a> reference, and check the <a href="https://www.elastic.co/docs/reference/query-languages/esql/limitations">ES|QL limitations</a> page for the current boundaries. It's a technical preview because we want your feedback: If you search something that was never indexed and it surprises you, either positively or negatively, <a href="https://www.elastic.co/community">we'd love to hear about it</a>.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/full-text-search-unindexed-data</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/full-text-search-unindexed-data</guid>
    <category><![CDATA[ES|QL]]></category>
    <category><![CDATA[Mappings]]></category>
    <dc:creator><![CDATA[Kevin Corcoran,Ioana Tagirta]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt9a74e633d05585a8/6a730774b8c2e64c3ebe0fd4/image1.jpg" length="0" type="image/jpeg"/>
    <pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[One field, every modality: how Elasticsearch's semantic field indexes and searches images, audio, video and PDFs automatically]]></title>
    <description><![CDATA[The semantic field turns images, audio, video, PDFs and text into multimodal embeddings at ingest time. Describe a scene and find the matching image or use a video frame to surface related clips, all from one Elasticsearch field.]]></description>
    <content:encoded><![CDATA[<p>Multimodal search in Elasticsearch now works the same way text search does: define a field, index your content, and query. The <code>semantic</code> field generates embeddings automatically at ingest time for images, audio, video, and PDFs. Every modality lands in one shared vector space, so you can retrieve an image with a text description, match audio to a phrase, or find a video with a still frame, all from a single field. Available in Elasticsearch 9.5 and serverless as a tech preview.</p><h2>The palette takes shape: how multimodal search in Elasticsearch evolved from semantic_text</h2><p>The <code>semantic</code> field is a convergence of several complementary features we've introduced over the past couple of years, bringing them together to create a cohesive multimodal search experience. Each solved an important piece of the semantic search puzzle on its own; together they enable native multimodal search.</p><p>The first brushstroke was <code>semantic_text</code>. Before it, running semantic search meant manually configuring mappings, wiring up ingest pipelines with an ML model, manually chunking content, and generating query-time embeddings yourself. The <code>semantic_text</code> field folds all of that away: it performs inference automatically at ingest time, chunks long documents for you, and simplifies the queries you write against it. <a href="https://www.elastic.co/search-labs/blog/semantic-search-simplified-semantic-text">Introduced in Elasticsearch 8.15</a> and <a href="https://www.elastic.co/search-labs/blog/elasticsearch-semantic-text-ga">released as GA in Elasticsearch 8.18</a>, it has become the foundation for semantic search on the platform.</p><p>Next came <a href="https://www.elastic.co/search-labs/blog/jina-embeddings-v5-omni-all-media-one-index">the model to power multimodal search</a>. <code>jina-embeddings-v5-omni</code> is our family of multimodal embedding models, capable of embedding text, images, video, audio, and PDFs into a shared vector space. Because those embeddings are semantically compatible across modalities, you can store diverse media in a single index and query across all of it at once, such as retrieving an image via a text description or matching audio against a written phrase, all without maintaining a separate pipeline for each content type. For more detailed information about how these embeddings are generated, see the <a href="https://jina.ai/models/jina-embeddings-v5-omni-small/">model documentation</a>.</p><p>We added the <a href="https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-knn-query#query-vector-builders-parameters">embedding query vector builder</a> in Elasticsearch 9.4 to handle multimodal inputs at query time. Query vector builders are general-purpose tools you can use to convert input to a vector at query time as part of your request. For example, we have the <code>text_embedding</code> query vector builder for text-only models and input, and the <code>lookup</code> query vector builder for getting a vector from an existing document. The <code>embedding</code> query vector builder is a new type that works with multimodal models and accepts multimodal input, including text or base64-encoded binaries. This allows you to pose a query in whatever modality fits, and Elasticsearch generates the matching vector on the fly.</p><p>The final piece was multimodal ingest. The <code>semantic_text</code> field brought automatic embedding to text; the <code>semantic</code> field extends that same automatic experience to images, audio, video, and PDFs from ingest through query.</p><h2>Painting the picture: creating an index with the semantic field</h2><p>Let’s create an index with a <code>semantic</code> field. This is as simple as setting the field type to semantic and defining the inference endpoint you want to use:</p>PUT example-index
{
  "mappings": {
    "properties": {
      "my_semantic_field": {
        "type": "semantic",
        "inference_id": ".jina-embeddings-v5-omni-small"
      }
    }
  }
}<p>In this example, we use the .<code>jina-embeddings-v5-omni-small</code> inference endpoint. This is our built-in <code>jina-embeddings-v5-omni</code> inference service, and it is available in all environments with access to the <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">Elastic Inference Service</a> (EIS). This includes:</p><ul><li><p>Serverless.</p></li><li><p>Elastic Cloud Hosted (ECH).</p></li><li><p>Self-managed with <a href="https://www.elastic.co/docs/deploy-manage/cloud-connect">Cloud Connected Mode</a> (CCM).</p></li></ul><h3>Indexing images, audio, video and PDFs</h3><p>To index an image, provide an object with a <code>type</code> of <code>image</code> and a <code>value</code> containing the image as a base64-encoded <a href="https://developer.mozilla.org/en-US/docs/Web/URI/Reference/Schemes/data">data URL</a>:</p>PUT example-index/_doc/example_doc_1
{
  "my_semantic_field": {
    "type": "image",
    "value": "data:image/jpeg;base64,&lt;base64-encoded-image-bytes&gt;"
  }
}<p>Arrays of objects are also accepted, allowing you to index multiple images in a single field value:</p>PUT example-index/_doc/example_doc_2
{
  "my_semantic_field": [
    {
      "type": "image",
      "value": "data:image/jpeg;base64,&lt;base64-encoded-image-bytes&gt;"
    },
    {
      "type": "image",
      "value": "data:image/jpeg;base64,&lt;base64-encoded-image-bytes&gt;"
    }
  ]
}<p>The <code>semantic</code> field also supports text values, just like <code>semantic_text</code>. You can provide such values standalone or intermix them with image values:</p>PUT example-index/_doc/example_doc_3
{
  "my_semantic_field": "a cat on a windowsill"                                                                                                                                                                                                                }

PUT example-index/_doc/example_doc_4
{
  "my_semantic_field": [
    "a cat on a windowsill",
    {
      "type": "image",
      "value": "data:image/jpeg;base64,&lt;base64-encoded-image-bytes&gt;"
    },
    "a dog running in a park"
  ]
}<p>Text values are handled just like they are with <code>semantic_text</code>: long passages are chunked according to the <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-field-reference#semantic-params">chunking settings configured on either the inference service or field mapping</a>. Multimodal values, such as images, are not chunked. Each multimodal value is represented as one chunk.</p><p>Other modalities are supported as well. Change the type value to match your content’s modality. Currently we support:</p><ul><li><p><code>image</code></p></li><li><p><code>audio</code></p></li><li><p><code>video</code></p></li><li><p><code>pdf</code></p></li></ul><p>For example, to index a video, the request would look like:</p>PUT example-index/_doc/example_doc_5
{
  "my_semantic_field": {
    "type": "video",
    "value": "data:video/mp4;base64,&lt;base64-encoded-video-bytes&gt;"
  }
}<p></p><h3>Image search and cross-modal retrieval with a text query</h3><p>To find multimodal content using a text description, run a <code>match</code> query on the <code>semantic</code> field:</p>GET example-index/_search
{
  "query": {
    "match": {
      "my_semantic_field": "a cat on a windowsill"
    }
  }
}<p>Just like with <code>semantic_text</code>, Elasticsearch automatically generates an embedding for the query text using the inference endpoint associated with the field. That query embedding is used to return semantically similar matches.</p><p>This query pattern enables easy text-to-image search. Just index an image and use a <code>match</code> query to retrieve it via text description! It also works for any other modality: index the multimodal input and search by description to retrieve it.</p><h3>Querying with images, video, and other multimodal inputs</h3><p>We can also search using a multimodal input by using the <code>knn</code> query with an <code>embedding</code> query vector builder. For example, we can search using an image:</p>GET example-index/_search
{
  "query": {
    "knn": {
      "field": "my_semantic_field",
      "query_vector_builder": {
        "embedding": {
          "input": {
            "type": "image",
            "value": "data:image/jpeg;base64,&lt;base64-encoded-image-bytes&gt;"
          }
        }
      }
    }
  }
}<p>The <code>input</code> object format is the same as when providing an image to index: set the <code>type</code> to <code>image</code> and <code>value</code> to a base64-encoded data URL.</p><p>Similar to when querying by text description, Elasticsearch automatically generates an embedding for the query image using the inference endpoint associated with the field. That query embedding is used to return semantically similar matches.</p><p>Just like with indexing, other modalities are supported, but are limited to those supported by your inference endpoint. For example, a search using a video clip would look like:</p>GET example-index/_search
{
  "query": {
    "knn": {
      "field": "my_semantic_field",
      "query_vector_builder": {
        "embedding": {
          "input": {
            "type": "video",
            "value": "data:video/mp4;base64,&lt;base64-encoded-video-bytes&gt;"
          }
        }
      }
    }
  }
}<h2>Extending the composition: highlighting, retrievers, and other semantic field features</h2><p>The <code>semantic</code> field didn't start from a blank canvas. It's built on the same foundation as <code>semantic_text</code>, inheriting its behavior and its ergonomics, and extending them to multimodal content. In practice, that means nearly everything you already know about working with <code>semantic_text</code> carries over unchanged. If you've built with <code>semantic_text</code> before, the <code>semantic</code> field will feel immediately familiar.</p><p>Here’s a selection of the features that come along for the ride. See <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-field">the documentation</a> for a complete list.</p><h3>Highlighting the best-matching chunks</h3><p>If you index multiple values in a <code>semantic</code> field, you may want to know <em>which</em> value best matches the query. The <code>semantic</code> highlighter can be used to return the most relevant chunks as highlight fragments:</p>GET example-index/_search
{
  "query": {
    "match": {
      "my_semantic_field": "a cat on a windowsill"
    }
  },
  "highlight": {
    "fields": {
      "my_semantic_field": {
        "number_of_fragments": 2,
        "order": "score"
      }
    }
  }
}<p>Setting <code>order</code> to <code>score</code> returns the fragments ranked by relevance, while <code>number_of_fragments</code> caps how many chunks come back. The response looks like:</p>{
  "hits": {
    "hits": [
      {
        "_index": "example-index",
        "_id": "example_doc_4",
        "_source": {...},
        "highlight": {
          "my_semantic_field": [
            "a cat on a windowsill",
            "data:image/jpeg;base64,&lt;base64-encoded-image-bytes&gt;"
          ]
        }
      }
    ]
  }
}<p>Note how highlighted multimodal values are represented using their data URLs.</p><h3>Controlling vector quantisation with index options</h3><p>The <code>semantic</code> field stores its embeddings in an underlying vector field, and <code>index_options</code> lets you control how that vector field is indexed. For example, choosing a non-default quantization strategy:</p>PUT example-index
{
  "mappings": {
    "properties": {
      "my_semantic_field": {
        "type": "semantic",
        "inference_id": ".jina-embeddings-v5-omni-small",
        "index_options": {
          "dense_vector": {
            "type": "int8_hnsw"
          }
        }
      }
    }
  }
}<h3>Multi-field retrievers</h3><p>The <code>semantic</code> field participates in the <a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/retrievers">multi-field query format</a> supported by the <code>linear</code> and <code>rrf</code> retrievers. Rather than hand-writing an inner retriever per field, you supply a single <code>query</code> and a list of <code>fields</code>, mixing lexical fields and semantic fields freely:</p>GET example-index/_search
{
  "retriever": {
    "linear": {
      "query": "a cat on a windowsill",
      "fields": ["title", "my_semantic_field"],
      "normalizer": "minmax"
    }
  }
}<p>The retriever automatically separates lexical fields from semantic fields, queries each group, and normalizes the results so that each group contributes equally to the final ranking, preventing lexical matches from drowning out semantic ones.</p><h3>Cross-cluster search</h3><p>The <code>semantic</code> field supports <a href="https://www.elastic.co/docs/solutions/search/cross-cluster-search">cross-cluster search (CCS)</a>, enabling use of the field in large, multi-cluster deployments. Simply list the indices to query using the standard <code>&lt;cluster&gt;:&lt;index&gt;</code> format:</p>GET example-index,remote-cluster:remote-index/_search
{
  "query": {
    "match": {
      "my_semantic_field": "a cat on a windowsill"
    }
  }
}<p>The fields queried across indices and clusters can use a mix of different inference endpoints that produce different query embeddings. The search request will automatically apply the proper query embedding to each individual field queried.</p><h2>Off the easel, into the world: optimising multimodal embeddings for production</h2><p>When you move multimodal search from experiment to production, the size of your multimodal inputs becomes a practical concern. Multimodal data is supplied as base64-encoded data URLs, and that data is stored in the index. Those strings can grow large in a hurry: a single high-resolution file can balloon into several megabytes of encoded text, which has several side effects:</p><ul><li><p>The index size on disk can increase significantly.</p></li><li><p>Requests and responses containing multimodal data are larger, increasing transmission time and ingress/egress costs.</p></li><li><p>Inference on larger multimodal inputs is slower.</p></li></ul><p>The good news is that you don’t need that much fidelity. Multimodal embedding models reduce each input to a compact representation before generating a vector anyway, so a smaller, lower-fidelity version of a multimodal input (such as a downscaled image or a lower-bitrate audio clip) produces a very similar embedding, and similar search quality, to its full-size original. This also applies to PDF input. PDFs are generally processed visually by multimodal models, so the quality only needs to be good enough to perform operations like image embedding and OCR. Long PDFs should be broken up into chunks of smaller inputs, so the embeddings generated more accurately represent each chunk. Feeding the model small inputs keeps your documents lean, trims index and response sizes, and speeds up ingestion, all without meaningfully affecting relevance. </p><p>Elasticsearch reinforces this practice with a guardrail: the <code>indices.inference.max_binary_input_size</code> cluster setting caps the size of each binary input, defaulting to 1 MB. Any individual value that exceeds the limit is rejected with a clear error, so oversized inputs surface as an actionable problem at index time rather than as silent bloat. This setting is adjustable in self-hosted and ECH through the <a href="https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-cluster-put-settings">cluster settings API</a>. It is not adjustable in our serverless offering, where 1 MB is the hard limit for binary sizes.</p><p>When possible, it is also advised to use <a href="https://www.elastic.co/docs/reference/elasticsearch/rest-apis/retrieve-selected-fields#source-filtering">source filtering</a> to exclude <code>semantic</code> fields from responses. For example:</p>GET example-index/_search
{ 
  "_source": {
    "excludes": ["my_semantic_field"]
  },
  "query": {
    "match": {
      "my_semantic_field": "a cat on a windowsill"
    }
  }
}<p>This makes responses smaller, more performant, and easier to parse because multimodal data is not returned with each.</p><h2>Try out the semantic field</h2><p>The <code>semantic</code> field is available in Elasticsearch 9.5 and Serverless. <a href="https://cloud.elastic.co/registration?onboarding_token=search&amp;cta=cloudregistration&amp;tech=trial&amp;plcmt=cross%20module&amp;pg=search-labs">Start a free trial</a> and try it out today.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/semantic-field-multimodal-search-elasticsearch</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/semantic-field-multimodal-search-elasticsearch</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[Index Data]]></category>
    <category><![CDATA[Jina AI]]></category>
    <category><![CDATA[Mappings]]></category>
    <dc:creator><![CDATA[Mike Pellegrini]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltac28de8857eafbc4/6a6f090aca9a724b3c614914/image1.png" length="0" type="image/png"/>
    <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>