Create and manage datasets for ES|QL Data Federation

Create and manage ES|QL Data Federation datasets in Kibana or with the /_query/dataset API. Before creating a dataset, connect a data source and review how to define a dataset.

Warning

This feature is experimental. It is not intended for production use and there are no guarantees around performance, scale, or stability in this release.

In Kibana, you create and manage datasets from the Datasets tab under Data management > ES|QL Data Federation.

The Datasets tab lists each dataset with its:

  • Data source and data source type
  • Resource
  • Description

From this tab you can search your datasets, filter by data source, add a new one, and edit or delete an existing one.

Click Add dataset to open a flyout where you define the dataset:

  • Data source: the connected data source to read through.
  • Name: a unique name for use in queries. Names must be lowercase and cannot begin with -, _, or +. A dataset cannot share a name with any existing index, data stream, alias, or view.
  • Description: an optional description (up to 1,000 characters).
  • Resource: the URI and glob pattern that selects the files to read. Refer to resource patterns for the pattern language.
  • Format: the file format. This selection is required in the Kibana UI. The API can omit settings.format when the resource pattern implies exactly one format. Extensionless or mixed patterns require format. Refer to supported file formats.

To configure how the format is read, expand Advanced settings. Refer to dataset settings.

To customize the inferred schema, rename columns, or override field types, declare a schema explicitly. Schema customization is not available in the UI.

Datasets are managed under the /_query/dataset endpoint. All dataset operations require the index manage privilege on the dataset name, or a fine-grained dataset privilege. Refer to manage credentials and privileges for details.

Operation Endpoint API reference
Create or update PUT /_query/dataset/{name} Create or update an ES|QL dataset
Get GET /_query/dataset/{name} Get ES|QL datasets
List all GET /_query/dataset Get ES|QL datasets
Delete DELETE /_query/dataset/{name} Delete ES|QL datasets

PUT /_query/dataset/{name} creates a new dataset or replaces an existing one entirely.

A dataset cannot have the same name as an existing index, data stream, alias, or view, because dataset names share the same namespace. Dataset names must be lowercase and cannot begin with -, _, or +.

The optional description can be at most 1,000 characters long.

For restrictions on S3 bucket names, aliases, and Amazon Resource Names (ARNs), refer to Amazon S3 resource restrictions.

				PUT /_query/dataset/access_logs
					{
  "data_source": "prod_s3_logs",
  "resource": "s3://logs-bucket/access/**/*.parquet",
  "description": "Production access logs",
  "settings": {
    "partition_detection": "hive"
  }
}
		
curl -X PUT "${ELASTICSEARCH_URL}/_query/dataset/access_logs" \
  -H "Authorization: ApiKey ${API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
  "data_source": "prod_s3_logs",
  "resource": "s3://logs-bucket/access/**/*.parquet",
  "description": "Production access logs",
  "settings": {
    "partition_detection": "hive"
  }
}'
		

Elasticsearch validates setting values when you register the dataset, rather than waiting until the first query. A value that isn't supported, such as a multi-character delimiter, an unknown encoding, or a segment_size below the minimum, returns a 400 error that identifies the setting.

Note

Datasets registered before this validation was introduced continue to work without being revalidated. Replacing one of these datasets triggers validation, so correct any unsupported values in the replacement request.

After creating a dataset, verify the inferred schema by checking its field mappings.

GET /_query/dataset/{name} retrieves a dataset by name.

				GET /_query/dataset/access_logs
		
curl -X GET "${ELASTICSEARCH_URL}/_query/dataset/access_logs" \
  -H "Authorization: ApiKey ${API_KEY}"
		

GET /_query/dataset returns all registered datasets.

				GET /_query/dataset
		
curl -X GET "${ELASTICSEARCH_URL}/_query/dataset" \
  -H "Authorization: ApiKey ${API_KEY}"
		

DELETE /_query/dataset/{name} deletes a dataset by name.

				DELETE /_query/dataset/access_logs
		
curl -X DELETE "${ELASTICSEARCH_URL}/_query/dataset/access_logs" \
  -H "Authorization: ApiKey ${API_KEY}"
		

After creating a dataset, query it with ES|QL. To change how the dataset selects or interprets files, return to define a dataset.