Dataset settings reference for ES|QL Data Federation

Dataset settings control how the files in a dataset are discovered, parsed, and reconciled. Add settings to the settings object when you create or update a dataset.

Warning

This feature is experimental. It is not intended for production use and there are no guarantees around performance, scale, or stability in this release.

The following settings apply to every file-based data source, unless an entry lists specific formats.

These settings select the format reader and filter the objects that wildcard discovery returns.

format

The file format reader used for the dataset.

  • Default: Inferred when the resource pattern implies exactly one format. Required otherwise.
  • Valid values: parquet, csv, tsv, ndjson, or auto to infer the format from the resource pattern

Set format for extensionless resources and for patterns that match more than one format. An explicit format sends objects with unrecognized extensions through this reader, but objects that map to a different registered format are rejected.

file_exclusions

Patterns that name objects to drop from wildcard discovery.

  • Default: ["**/_*", "**/.*", "**/_temporary/**", "**/_delta_log/**"]
  • Valid values: An array of patterns in the resource pattern language, or [] to turn off exclusion
  • Related: resource

Setting file_exclusions replaces the default list. To keep the defaults, include them in your list. For matching rules and examples, refer to exclude non-data objects.

These settings control how partition columns are derived from object paths.

partition_detection

How partition columns are derived from directory names.

  • Default: auto
  • Valid values:
    • auto: Reads Hive key=value directory names. When partition_path is set, uses that template for paths that don't use key=value.
    • hive: Reads Hive key=value directory names only.
    • template: Names partition columns from partition_path.
    • none: Turns off partition detection.
  • Requires: partition_path when set to template
  • Conflicts with:
    • partition_path when set to hive or none
    • partition_spec when set to none
  • Related: partition_path, partition_spec, partition_sample_size

partition_path

A template that names partition columns for paths that don't use key=value directories.

  • Default: None
  • Valid values: A path template that uses {column} placeholders, for example {year}/{month}
  • Conflicts with: partition_detection set to hive or none
  • Related: partition_detection, partition_spec

Each placeholder labels one path segment. For example, {year}/{month} extracts year and month columns from a two-level path. The default partition_detection of auto reads Hive directory names first and uses the template for other paths. For placeholder syntax, refer to define partition paths.

partition_spec

Maps file columns to partition keys, so that filters on those columns can skip folders.

  • Default: None
  • Valid values: A comma-separated list of bindings, each in one of these forms:
    • [key=]transform(column[, unit]): A temporal or identity transform. transform is identity, year, month, day, or hour. unit is epoch_second or epoch_millis, and applies only to temporal transforms. The default unit is epoch_millis. Unit names follow the date format names.
    • key=column: Maps a column to a differently named key.
    • column: Maps a column to the key with the same name.
  • Requires: Each key to be a {name} placeholder in partition_path, when partition_path is set
  • Conflicts with: partition_detection set to none
  • Related: partition_detection, partition_path

For syntax, examples, and how folders are skipped, refer to Skip folders with file column filters.

partition_sample_size

The number of file paths read to infer partition columns and their types.

  • Default: 1000
  • Valid values: An integer from 1 through 10000000
  • Related: partition_detection, schema_resolution

The sample is the first paths in listing order, not a random selection. Raise the value when some partition values first appear later in the listing, so that those values get a column.

The sample applies only to a query that reads no rows, and only when the listing covers the whole dataset in the store's own order. In the following cases, every file is listed and the sample size has no effect:

  • The query reads rows.
  • The query filters on a partition column or on _file.*.
  • The dataset sets file_sort_by or file_order to a value other than the default.
  • The dataset uses union_by_name or strict schema resolution.

These settings control how schemas are combined when a dataset spans multiple files. For concepts and examples, refer to schema inference.

schema_resolution

The strategy for reconciling schemas across multiple files.

  • Default:
    • first_file_wins
    • union_by_name
  • Valid values: first_file_wins, union_by_name, strict
  • Related: file_sort_by, file_order

To compare strategies and learn how existing datasets resolve a missing value, refer to choose a schema resolution strategy.

file_sort_by

The value used to order files when choosing which file supplies the schema.

  • Default: list
  • Valid values: list, name, mtime
  • Requires: An effective schema_resolution of first_file_wins
  • Related: file_order

For how each value orders files, refer to control which file supplies the schema.

file_order

The sort direction for file_sort_by.

  • Default: asc
  • Valid values: asc, desc
  • Requires: An effective schema_resolution of first_file_wins
  • Related: file_sort_by

The direction also applies when file_sort_by is list. In that case, desc reverses the declaration or listing order.

These settings control how malformed rows are handled and how many are tolerated before the query fails.

error_mode

How malformed rows are handled.

  • Default: fail_fast
  • Valid values:
    • fail_fast: Fails the query at the first malformed row.
    • skip_row: Drops each malformed row.
    • null_field: Replaces a value that fails to parse with null and keeps the row.
  • Related: max_errors, max_error_ratio

max_errors

The maximum number of malformed rows allowed before the query fails.

  • Default: Unlimited
  • Valid values: A non-negative integer
  • Requires: An explicit error_mode of skip_row or null_field
  • Conflicts with: error_mode set to fail_fast
  • Related: max_error_ratio

max_error_ratio

The maximum fraction of malformed rows allowed before the query fails.

  • Default: 0.0, which applies no ratio limit
  • Valid values: A number from 0.0 through 1.0
  • Requires: An explicit error_mode of skip_row or null_field
  • Conflicts with: error_mode set to fail_fast
  • Related: max_errors

These settings control how files are divided into splits that nodes read in parallel.

target_split_size

The target size of each unit of work that a file is divided into for parallel reading across nodes.

  • Default: 64 MiB (64mb)
  • Valid values: A positive byte size, for example 32mb
  • Related: split_probe_window, max_split_probes

Files larger than the target are cut into several splits. Files smaller than the target are read as a single split. Lower the value for more parallelism over a few large files. Raise it to reduce planning work when a dataset contains a very large number of bytes.

split_probe_window

The number of bytes that each record-boundary search can read while files are split.

  • Formats: NDJSON, and CSV and TSV without quoting or escaping
  • Default: 256 KiB (256kb)
  • Valid values: A positive byte size. The product of max_split_probes and split_probe_window can't exceed 4 GiB.
  • Related: max_split_probes, target_split_size

If a dataset's records are longer than this value, its files are cut into fewer splits than target_split_size requests, which reduces parallelism. Raise the value for datasets with long records. If the pair is rejected, lower one of the two probe settings.

max_split_probes

The maximum number of record-boundary searches that a query can perform, which limits how many splits its files are cut into.

  • Formats: NDJSON, and CSV and TSV without quoting or escaping
  • Default: 1000
  • Valid values: An integer from 1 through 10000. The product of max_split_probes and split_probe_window can't exceed 4 GiB.
  • Related: split_probe_window, target_split_size

When a scan needs more splits than this value allows, the scan uses a larger split size than target_split_size requests. Raise the value to get the requested split size on a very large scan.

The following settings apply only to data sources that use a specific storage provider.

These settings apply to datasets whose data source uses Amazon S3 or an S3-compatible store.

region

The AWS region used for the S3 client, for example eu-central-1.

  • Default: Auto-detected
  • Valid values: A non-empty AWS region name
  • Related: endpoint and sts_region on the data source

Omit region for standard AWS S3. Set it when the data source uses a custom endpoint, such as MinIO or Scaleway, to skip region discovery on the first request.

The following settings apply to CSV and TSV files.

These settings cover the field separator, quoting style, header handling, and null tokens that most CSV and TSV files need.

delimiter

The field separator.

  • Default: , for CSV, \t for TSV
  • Valid values:
    • A single ASCII character other than a line feed or carriage return. Write a tab as \t and a backslash as \\.
    • Multi-character values are rejected when you create or update the dataset.
  • Conflicts with: The quote character when quoting is on, and the escape character when escaping is on
  • Related: quote, escape

mode

A preset that sets quoting and escaping together.

  • Default: quoted for CSV, plain for TSV
  • Valid values:
    • quoted: Fields can be wrapped in quotes. An embedded quote is doubled, and a backslash escapes characters inside a quoted field.
    • escaped: No quoting. A backslash escapes special characters, and \N reads as null.
    • plain: No quoting or escaping. Every byte is literal, so a field can't contain the delimiter or a newline.
  • Conflicts with: An explicit quote when set to escaped
  • Related: quote, escape

An explicit quote or escape value overrides the preset.

header_row

Whether the first record that isn't blank or a comment names the columns.

  • Default: true
  • Valid values: true, false
  • Related: skip_rows, column_prefix

header_row is applied after skip_rows.

skip_rows

The number of leading content records to discard from each file.

  • Default: 0
  • Valid values: An integer from 0 through 1000
  • Related: header_row, comment

Records are discarded after decompression and before header_row is applied. Blank lines and comment lines don't count toward skip_rows. For example, read a file that starts with two prose lines and then the header state,ip,user_agent with "skip_rows": 2 and "header_row": true. A preamble of comment lines is skipped through comment without setting skip_rows.

null_value

The token that reads as null.

  • Default: None. No token reads as null.
  • Valid values: A string, for example NULL, NA, or \N.

An empty string "" makes empty fields read as null.

encoding

The file's character encoding.

  • Default: UTF-8
  • Valid values: A character set name, for example ISO-8859-1

These settings tune schema sampling, quoting characters, column naming, value parsing, and field size limits.

schema_sample_size

The number of rows sampled to infer the schema.

  • Default: 20000
  • Valid values:
    • An integer from 1 through 20000
    • An integer from 1 through 1000

The sample determines whether sparse or late-appearing fields get a column. To learn how schemas are inferred, refer to schema inference.

quote

The quote character.

  • Default: " for CSV. Quoting is off for TSV.
  • Valid values:
    • A single ASCII character other than a line feed or carriage return. Write a tab as \t and a backslash as \\.
    • none to turn off quoting
    • Multi-character values are rejected when you create or update the dataset.
  • Conflicts with:
    • The delimiter character
    • The escape character, when escaping is on
    • mode set to escaped
  • Related: mode, escape, delimiter

An explicit value overrides the mode preset.

escape

The escape character.

  • Default: \ for CSV. Escaping is off for TSV.
  • Valid values:
    • A single ASCII character other than a line feed or carriage return. Write a tab as \t and a backslash as \\.
    • none to turn off escaping
    • Multi-character values are rejected when you create or update the dataset.
  • Conflicts with:
    • The delimiter character
    • The quote character, when quoting is on
  • Related: mode, quote, delimiter

An explicit value overrides the mode preset.

comment

The prefix that marks a line as a comment to skip.

  • Default: //
  • Valid values: A string

column_prefix

The prefix for generated column names when header_row is false.

  • Default: col
  • Valid values: A string
  • Related: header_row

Each name ends with a counter that starts at 0, for example col0, col1, col2. An empty prefix produces numeric column names, which must be quoted with backticks in ES|QL.

datetime_format

The pattern used to parse date and time values.

  • Default: ISO 8601 or epoch milliseconds
  • Valid values: A date format pattern or built-in format name. Combine formats with ||.

trim_spaces

Whether to remove surrounding ASCII whitespace from string field values.

  • Default: false
  • Valid values: true, false

Typed values, such as numbers and dates, tolerate surrounding whitespace regardless of this setting.

multi_value_syntax

Whether bracketed multi-values are recognized.

  • Default: none
  • Valid values:
    • none: Reads brackets as literal characters.
    • brackets: Reads a field such as [a,b,c] as a multi-value.
  • Requires: Quoting when set to brackets. Without a mode, brackets selects quoted.
  • Conflicts with: mode set to escaped or plain, or quote set to none, when set to brackets

max_field_size

The maximum size of a single field, in bytes.

  • Default: 10 MiB (10485760)
  • Valid values: An integer number of bytes. 0 removes the limit.

The following settings apply to NDJSON files.

This setting controls how much of each file is sampled to infer the schema.

schema_sample_size

The number of lines sampled to infer the schema.

  • Default: 20000
  • Valid values:
    • An integer from 1 through 20000
    • An integer from 1 through 1000

The sample determines whether sparse or late-appearing fields get a column. To learn how schemas are inferred, refer to schema inference.

These settings tune parallel reading, date parsing, and schema size limits for NDJSON files.

segment_size

The unit that a file is divided into for parallel reading.

  • Default: 4 MiB (4mb)
  • Valid values: A byte size of at least 64 KiB (64kb)

datetime_format

The pattern used to infer and parse date and time values.

  • Default: strict_date_optional_time
  • Valid values: A date format pattern or built-in format name. Combine formats with ||.

schema_max_fields

The maximum number of fields that schema inference can create from a file.

  • Default: 1000, or the value of the esql.external.schema_max_fields cluster setting
  • Valid values: An integer from 1 through 100000

Objects count as fields, as well as leaf fields, and each segment of a dotted key counts as a field. If a file's inferred schema exceeds the limit, the query fails.

Parquet is self-describing and has no format-specific dataset settings.