ES|QL DENSE_VECTOR command
The DENSE_VECTOR command generates vector embeddings from one or more text
or keyword columns. It embeds each row independently through an inference
endpoint and appends the resulting dense_vector columns to the query output.
You can embed indexed or computed text, or base64-encoded image data.
DENSE_VECTOR <input_column> [, <input_column>, ...] [WITH { <options> }]
DENSE_VECTOR <output_column> = <input_column> [WITH { <options> }]
DENSE_VECTOR suffix = "<suffix>" ON <input_column> [, <input_column>, ...] [WITH { <options> }]
input_column- (Required) One or more comma-separated
textorkeywordcolumns to embed. For image input, the column must contain a base64-encoded image data URI. A non-string column is rejected before execution. If a value isnull, the generated vector is alsonull. output_column- (Optional) The name of the generated column. You can specify an output column name only when embedding one input column. If the name matches an existing column, the generated column replaces it.
suffix- (Optional) A quoted string appended to each input column name. For example,
DENSE_VECTOR suffix = "_dv" ON title, bodyproducestitle_dvandbody_dv. You can use a suffix with one or more input columns. Without a naming clause, output columns use the name<input_column>_dense_vector.
inference_id- (Optional) The ID of the
inference endpoint
used to embed the input. If omitted,
DENSE_VECTORselects a default text embedding endpoint. See Inference endpoints. type- (Optional) The input modality. Accepts
text(default) orimage. Animageinput is a base64 data URI embedded through a multimodal endpoint. timeout- (Optional) Timeout for the inference request (for example,
"30s","1m"). If not specified, the default inference timeout applies.
DENSE_VECTOR adds one dense_vector output column for each input column and
preserves the other columns in the result. You can embed columns loaded from an
index or columns computed earlier in the query, for example with EVAL.
Use the generated vectors with vector similarity functions such as
V_COSINE
or commands such as MMR.
Use DENSE_VECTOR to embed values that vary between rows. To embed a single
constant value, such as query text, use the
TEXT_EMBEDDING
function.
Learn more about using ES|QL for search use cases.
DENSE_VECTOR uses an
inference endpoint
to embed its input. The endpoint must support the selected input type:
- For text input, the endpoint must use the
text_embeddingor multimodalembeddingtask type. If you omitinference_id,DENSE_VECTORusesesql.command.dense_vector.default_inference_idwhen configured. Otherwise, it selects the first available built-in text embedding endpoint: the.jina-embeddings-v5-text-smallendpoint on the Elastic Inference Service (EIS), then the.multilingual-e5-small-elasticsearchendpoint, which runs on ML nodes. - For image input, the endpoint must use the multimodal
embeddingtask type. No default image endpoint is available, so you must specifyinference_id.
If no endpoint resolves, for example on a deployment that has neither built-in
endpoint, the query fails with an error naming each candidate and why it was
rejected. Specify inference_id to select an endpoint explicitly.
Embeddings produced by different models generally do not share the same vector
space. Specify inference_id when generated vectors must remain comparable
across queries or deployments.
In a cross-cluster query
or a cross-project query,
DENSE_VECTOR runs on the cluster or project that receives the query. The
inference endpoint must exist there, even when the documents come from a remote
cluster or a linked project.
Use the following dynamic cluster settings to control DENSE_VECTOR resource
usage and availability:
| Setting | Default | Purpose |
|---|---|---|
esql.command.dense_vector.enabled |
true |
Controls whether the command is available. |
esql.command.dense_vector.limit |
1000 |
Sets the maximum number of input rows processed by the command. |
esql.command.dense_vector.batch_size |
20 |
Sets the maximum number of inputs combined in one inference request. Accepts values from 1 to 1000. |
esql.command.dense_vector.default_inference_id |
Not set | Selects a default endpoint when a query omits inference_id. When not set, the command selects a built-in endpoint as described in Requirements. |
For a multivalued input column, DENSE_VECTOR embeds only the first value and
reports a warning for the discarded values. The first value depends on how the
column is loaded. To select the value explicitly, reduce the column first with a
function such as
MV_FIRST.
The following examples use columns from a books index, unless noted otherwise.
Embed a text column and add an output column named <input_column>_dense_vector:
FROM books
| WHERE book_no IN ("7480", "8605")
| DENSE_VECTOR title WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| KEEP book_no, title, title_dense_vector
| book_no:keyword | title:text | title_dense_vector:dense_vector |
|---|---|---|
| 7480 | The Hobbit | [50.0, 49.0, 48.0] |
| 8605 | Dead Souls | [45.0, 49.0, 53.0] |
List multiple input columns to embed them with one command:
FROM books
| WHERE book_no IN ("7480", "8605")
| DENSE_VECTOR title, publisher WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| KEEP book_no, title, publisher, title_dense_vector, publisher_dense_vector
| book_no:keyword | title:text | publisher:text | title_dense_vector:dense_vector | publisher_dense_vector:dense_vector |
|---|---|---|---|---|
| 7480 | The Hobbit | Mariner Books | [50.0, 49.0, 48.0] | [45.0, 49.0, 53.0] |
| 8605 | Dead Souls | Vintage | [45.0, 49.0, 53.0] | [50.0, 49.0, 50.0] |
Use output_column = input_column to replace the default output name:
FROM books
| WHERE book_no IN ("7480", "8605")
| DENSE_VECTOR book_vector = title WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| KEEP book_no, title, book_vector
| book_no:keyword | title:text | book_vector:dense_vector |
|---|---|---|
| 7480 | The Hobbit | [50.0, 49.0, 48.0] |
| 8605 | Dead Souls | [45.0, 49.0, 53.0] |
Use suffix = "..." ON to replace the default _dense_vector suffix for
each listed column:
FROM books
| WHERE book_no IN ("7480", "8605")
| DENSE_VECTOR suffix = "_vec" ON title, publisher WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| KEEP book_no, title_vec, publisher_vec
| book_no:keyword | title_vec:dense_vector | publisher_vec:dense_vector |
|---|---|---|
| 7480 | [50.0, 49.0, 48.0] | [45.0, 49.0, 53.0] |
| 8605 | [45.0, 49.0, 53.0] | [50.0, 49.0, 50.0] |
Embed a text or keyword column created earlier in the query. In this example,
EVAL creates the value that DENSE_VECTOR embeds:
FROM books
| WHERE book_no IN ("7480", "8605")
| EVAL description = CONCAT(title, " (", publisher, ")")
| DENSE_VECTOR description WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| KEEP book_no, description, description_dense_vector
| book_no:keyword | description:keyword | description_dense_vector:dense_vector |
|---|---|---|
| 7480 | The Hobbit (Mariner Books) | [45.0, 49.0, 49.0] |
| 8605 | Dead Souls (Vintage) | [53.0, 48.0, 53.0] |
Pass the generated column to MMR
to remove near-duplicate rows. MMR can use the runtime dense_vector, so this
workflow does not require an indexed vector field:
FROM books
| DENSE_VECTOR title WITH { "inference_id" : "test_dense_inference" }
| SORT book_no
| LIMIT 10
| MMR ON title_dense_vector LIMIT 3
| KEEP book_no, title
| book_no:keyword | title:text |
|---|---|
| 1211 | The brothers Karamazov |
| 1502 | Selected Passages from Correspondence with Friends |
| 1937 | The Best Short Stories of Dostoevsky (Modern Library) |
Set "type": "image" to embed a base64-encoded image data URI through a
multimodal endpoint. This example uses ROW because the input is a literal data
URI rather than an indexed column:
ROW input = "data:image/jpeg;base64,V2hvIGlzIFZpY3RvciBIdWdvPw=="
| DENSE_VECTOR input WITH { "inference_id" : "test_embedding_inference", "type" : "image" }
| input:keyword | input_dense_vector:dense_vector |
|---|---|
| data:image/jpeg;base64,V2hvIGlzIFZpY3RvciBIdWdvPw== | [-56.0, -50.0, -48.0] |