stack es ml start-trained-model-deployment cli command
Auth required
elastic stack es ml start-trained-model-deployment \
--model-id <model-id> \
[options]
Start a trained model deployment.
Behaviour flags:
--dry-run — validate all inputs and exit without performing any action
--model-idstringrequired- The unique identifier of the trained model. Currently, only PyTorch models are supported.
--cache-sizestring- The inference cache size (in memory outside the JVM heap) per node for the model.
The default value is the same size as the
model_size_bytes. To disable the cache,0bcan be provided. --deployment-idstring- A unique identifier for the deployment of the model.
--number-of-allocationsnumber- The number of model allocations on each node where the model is deployed. All allocations on a node share the same copy of the model in memory but use a separate set of threads to evaluate the model. Increasing this value generally increases the throughput. If this setting is greater than the number of hardware threads it will automatically be changed to a value less than the number of hardware threads. If adaptive_allocations is enabled, do not set this value, because it’s automatically set.
--priorityenum-
The deployment priority
Values: normal, low
--queue-capacitynumber- Specifies the number of inference requests that are allowed in the queue. After the number of requests exceeds this value, new requests are rejected with a 429 error.
--threads-per-allocationnumber- Sets the number of threads used by each model allocation during inference. This generally increases the inference speed. The inference process is a compute-bound process; any number greater than the number of available hardware threads on the machine does not increase the inference speed. If this setting is greater than the number of hardware threads it will automatically be changed to a value less than the number of hardware threads.
--timeoutstring- Specifies the amount of time to wait for the model to deploy.
--wait-forenum-
Specifies the allocation status to wait for before returning.
Values: started, starting, fully_allocated
--adaptive-allocationsstring- Adaptive allocations configuration. When enabled, the number of allocations is set based on the current load. If adaptive_allocations is enabled, do not set the number of allocations manually.
--error-trace- When set to
trueElasticsearch will include the full stack trace of errors when they occur. --filter-pathstring-
Comma-separated list of filters in dot notation which reduce the response returned by Elasticsearch.
Repeatable: pass
--filter-pathmultiple times to supply more than one value --human- When set to
truewill return statistics in a format suitable for humans. For example"exists_time": "1h"for humans and"exists_time_in_millis": 3600000for computers. When disabled the human readable values will be omitted. This makes sense for responses being consumed only by machines. --pretty- If set to
truethe returned JSON will be "pretty-formatted". Only use this option for debugging only. --input-filestring- path to a JSON file to use as command input
--dry-run- validate all inputs and exit without performing any action (preview changes without applying them)
--json-
output as JSON