ES|QL DENSE_VECTOR command

The DENSE_VECTOR command generates vector embeddings from one or more text or keyword columns. It embeds each row independently through an inference endpoint and appends the resulting dense_vector columns to the query output. You can embed indexed or computed text, or base64-encoded image data.

DENSE_VECTOR <input_column> [, <input_column>, ...] [WITH { <options> }]
DENSE_VECTOR <output_column> = <input_column> [WITH { <options> }]
DENSE_VECTOR suffix = "<suffix>" ON <input_column> [, <input_column>, ...] [WITH { <options> }]
		
input_column
(Required) One or more comma-separated text or keyword columns to embed. For image input, the column must contain a base64-encoded image data URI. A non-string column is rejected before execution. If a value is null, the generated vector is also null.
output_column
(Optional) The name of the generated column. You can specify an output column name only when embedding one input column. If the name matches an existing column, the generated column replaces it.
suffix
(Optional) A quoted string appended to each input column name. For example, DENSE_VECTOR suffix = "_dv" ON title, body produces title_dv and body_dv. You can use a suffix with one or more input columns. Without a naming clause, output columns use the name <input_column>_dense_vector.
inference_id
(Optional) The ID of the inference endpoint used to embed the input. If omitted, DENSE_VECTOR selects a default text embedding endpoint. See Inference endpoints.
type
(Optional) The input modality. Accepts text (default) or image. An image input is a base64 data URI embedded through a multimodal endpoint.
timeout
(Optional) Timeout for the inference request (for example, "30s", "1m"). If not specified, the default inference timeout applies.

DENSE_VECTOR adds one dense_vector output column for each input column and preserves the other columns in the result. You can embed columns loaded from an index or columns computed earlier in the query, for example with EVAL.

Use the generated vectors with vector similarity functions such as V_COSINE or commands such as MMR.

Tip

Use DENSE_VECTOR to embed values that vary between rows. To embed a single constant value, such as query text, use the TEXT_EMBEDDING function.

Learn more about using ES|QL for search use cases.

DENSE_VECTOR uses an inference endpoint to embed its input. The endpoint must support the selected input type:

  • For text input, the endpoint must use the text_embedding or multimodal embedding task type. If you omit inference_id, DENSE_VECTOR uses esql.command.dense_vector.default_inference_id when configured. Otherwise, it selects the first available built-in text embedding endpoint: the .jina-embeddings-v5-text-small endpoint on the Elastic Inference Service (EIS), then the .multilingual-e5-small-elasticsearch endpoint, which runs on ML nodes.
  • For image input, the endpoint must use the multimodal embedding task type. No default image endpoint is available, so you must specify inference_id.

If no endpoint resolves, for example on a deployment that has neither built-in endpoint, the query fails with an error naming each candidate and why it was rejected. Specify inference_id to select an endpoint explicitly.

Embeddings produced by different models generally do not share the same vector space. Specify inference_id when generated vectors must remain comparable across queries or deployments.

In a cross-cluster query or a cross-project query, DENSE_VECTOR runs on the cluster or project that receives the query. The inference endpoint must exist there, even when the documents come from a remote cluster or a linked project.

Use the following dynamic cluster settings to control DENSE_VECTOR resource usage and availability:

Setting Default Purpose
esql.command.dense_vector.enabled true Controls whether the command is available.
esql.command.dense_vector.limit 1000 Sets the maximum number of input rows processed by the command.
esql.command.dense_vector.batch_size 20 Sets the maximum number of inputs combined in one inference request. Accepts values from 1 to 1000.
esql.command.dense_vector.default_inference_id Not set Selects a default endpoint when a query omits inference_id. When not set, the command selects a built-in endpoint as described in Requirements.

For a multivalued input column, DENSE_VECTOR embeds only the first value and reports a warning for the discarded values. The first value depends on how the column is loaded. To select the value explicitly, reduce the column first with a function such as MV_FIRST.

The following examples use columns from a books index, unless noted otherwise.

Embed a text column and add an output column named <input_column>_dense_vector:

FROM books
| WHERE book_no IN ("7480", "8605")
| DENSE_VECTOR title WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| KEEP book_no, title, title_dense_vector
		
book_no:keyword title:text title_dense_vector:dense_vector
7480 The Hobbit [50.0, 49.0, 48.0]
8605 Dead Souls [45.0, 49.0, 53.0]

List multiple input columns to embed them with one command:

FROM books
| WHERE book_no IN ("7480", "8605")
| DENSE_VECTOR title, publisher WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| KEEP book_no, title, publisher, title_dense_vector, publisher_dense_vector
		
book_no:keyword title:text publisher:text title_dense_vector:dense_vector publisher_dense_vector:dense_vector
7480 The Hobbit Mariner Books [50.0, 49.0, 48.0] [45.0, 49.0, 53.0]
8605 Dead Souls Vintage [45.0, 49.0, 53.0] [50.0, 49.0, 50.0]

Use output_column = input_column to replace the default output name:

FROM books
| WHERE book_no IN ("7480", "8605")
| DENSE_VECTOR book_vector = title WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| KEEP book_no, title, book_vector
		
book_no:keyword title:text book_vector:dense_vector
7480 The Hobbit [50.0, 49.0, 48.0]
8605 Dead Souls [45.0, 49.0, 53.0]

Use suffix = "..." ON to replace the default _dense_vector suffix for each listed column:

FROM books
| WHERE book_no IN ("7480", "8605")
| DENSE_VECTOR suffix = "_vec" ON title, publisher WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| KEEP book_no, title_vec, publisher_vec
		
book_no:keyword title_vec:dense_vector publisher_vec:dense_vector
7480 [50.0, 49.0, 48.0] [45.0, 49.0, 53.0]
8605 [45.0, 49.0, 53.0] [50.0, 49.0, 50.0]

Embed a text or keyword column created earlier in the query. In this example, EVAL creates the value that DENSE_VECTOR embeds:

FROM books
| WHERE book_no IN ("7480", "8605")
| EVAL description = CONCAT(title, " (", publisher, ")")
| DENSE_VECTOR description WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| KEEP book_no, description, description_dense_vector
		
book_no:keyword description:keyword description_dense_vector:dense_vector
7480 The Hobbit (Mariner Books) [45.0, 49.0, 49.0]
8605 Dead Souls (Vintage) [53.0, 48.0, 53.0]

Pass the generated column to MMR to remove near-duplicate rows. MMR can use the runtime dense_vector, so this workflow does not require an indexed vector field:

FROM books
| DENSE_VECTOR title WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| LIMIT 10
| MMR ON title_dense_vector LIMIT 3
| KEEP book_no, title
		
book_no:keyword title:text
1211 The brothers Karamazov
1502 Selected Passages from Correspondence with Friends
1937 The Best Short Stories of Dostoevsky (Modern Library)

Set "type": "image" to embed a base64-encoded image data URI through a multimodal endpoint. This example uses ROW because the input is a literal data URI rather than an indexed column:

ROW input = "data:image/jpeg;base64,V2hvIGlzIFZpY3RvciBIdWdvPw=="
| DENSE_VECTOR input WITH { "inference_id" : "test_embedding_inference", "type" : "image" }
		
input:keyword input_dense_vector:dense_vector
data:image/jpeg;base64,V2hvIGlzIFZpY3RvciBIdWdvPw== [-56.0, -50.0, -48.0]