The chat completion inference API enables rich responses for chat completion tasks.
It only works with the chat_completion task type.
NOTE: The chat_completion task type supports both streaming and non-streaming.
The Chat completion inference API provides more comprehensive customization options through more fields and function calling support.
To determine whether a given inference service supports this task type, please see the page for that service.
These services support non-streaming chat completion inference:
- AI21
- Azure OpenAI
- Deepseek
- Elastic
- FireworksAI
- Groq
- Huggingface
- IBMWatsonX
- Llama
- Mistral
- NVIDIA
- OpenAI
- OpenShiftAI
Query parameters
-
Specifies the amount of time to wait for the inference request to complete.
Default value is
120s.
Body
Required
-
A list of objects representing the conversation. Requests should generally only add new messages from the user (role
user). The other message roles (assistant,system, ortool) should generally only be copied from the response to a previous completion request, such that the messages array is built up throughout a conversation. -
The ID of the model to use. By default, the model ID is set to the value included when creating the inference endpoint.
-
The upper bound limit for the number of tokens that can be generated for a completion request.
-
The reasoning configuration for the completion request. This controls the model's reasoning process in one of two ways:
- By specifying the model’s reasoning effort level with the
effortfield. - By enabling reasoning with default settings by setting
enabledfield totrue.
It also includes optional settings to control:
- The level of detail in the summary returned in the response with the
summaryfield. - Whether reasoning details are included in the response at all with the
excludefield.
Example (effort):
{ "reasoning": { "effort": "high", "summary": "concise", "exclude": false } }Example (enabled):
{ "reasoning": { "enabled": true, "summary": "concise", "exclude": false } }Currently supported only for
elasticprovider. - By specifying the model’s reasoning effort level with the
-
A sequence of strings to control when the model should stop generating additional tokens.
-
The sampling temperature to use.
- tool_choice
string | object Controls which tool is called by the model. String representation: One of
auto,none, orrequrired.autoallows the model to choose between calling tools and generating a message.nonecauses the model to not call any tools.requiredforces the model to call one or more tools. Example (object representation):{ "tool_choice": { "type": "function", "function": { "name": "get_current_weather" } } } -
A list of tools that the model can call. Example:
{ "tools": [ { "type": "function", "function": { "name": "get_price_of_item", "description": "Get the current price of an item", "parameters": { "type": "object", "properties": { "item": { "id": "12345" }, "unit": { "type": "currency" } } } } } ] } -
Nucleus sampling, an alternative to sampling with temperature.
POST _inference/chat_completion/openai-completion
{
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": "What is Elastic?"
}
]
}
resp = client.inference.non_streaming_chat_completion(
inference_id="openai-completion",
chat_completion_request={
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": "What is Elastic?"
}
]
},
)
const response = await client.inference.nonStreamingChatCompletion({
inference_id: "openai-completion",
chat_completion_request: {
model: "gpt-4o",
messages: [
{
role: "user",
content: "What is Elastic?",
},
],
},
});
response = client.inference.non_streaming_chat_completion(
inference_id: "openai-completion",
body: {
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": "What is Elastic?"
}
]
}
)
$resp = $client->inference()->nonStreamingChatCompletion([
"inference_id" => "openai-completion",
"body" => [
"model" => "gpt-4o",
"messages" => array(
[
"role" => "user",
"content" => "What is Elastic?",
],
),
],
]);
curl -X POST -H "Authorization: ApiKey $ELASTIC_API_KEY" -H "Content-Type: application/json" -d '{"model":"gpt-4o","messages":[{"role":"user","content":"What is Elastic?"}]}' "$ELASTICSEARCH_URL/_inference/chat_completion/openai-completion"
{
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": "What is Elastic?"
}
]
}
{
"messages": [
{
"role": "assistant",
"content": "Let's find out what the weather is",
"tool_calls": [
{
"id": "call_KcAjWtAww20AihPHphUh46Gd",
"type": "function",
"function": {
"name": "get_current_weather",
"arguments": "{\"location\":\"Boston, MA\"}"
}
}
]
},
{
"role": "tool",
"content": "The weather is cold",
"tool_call_id": "call_KcAjWtAww20AihPHphUh46Gd"
}
]
}
{
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "What's the price of a scarf?"
}
]
}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_current_price",
"description": "Get the current price of a item",
"parameters": {
"type": "object",
"properties": {
"item": {
"id": "123"
}
}
}
}
}
],
"tool_choice": {
"type": "function",
"function": {
"name": "get_current_price"
}
}
}
{
"messages": [{
"role": "user",
"content": [{
"type": "text",
"text": "Barber shaves all those, who do not shave themselves. Who shaves the barber?"
}
]
}, {
"role": "assistant",
"content": [{
"type": "text",
"text": "This is the barber paradox. Such a barber cannot logically exist."
}
],
"reasoning": "If the barber shaves himself, he should not; if he does not, he should.",
"reasoning_details": [{
"type": "reasoning.encrypted",
"data": "[REDACTED]"
}, {
"type": "reasoning.summary",
"summary": "Barber shaving himself creates contradiction"
}, {
"type": "reasoning.text",
"text": "If the barber shaves himself, he should not; if he does not, he should.",
"signature": "sig_123"
}
]
}, {
"role": "user",
"content": [{
"type": "text",
"text": "What if there are 2 barbers?"
}
]
}
],
"reasoning": {
"effort": "high",
"summary": "detailed",
"exclude": false
}
}
{
"messages": [{
"role": "user",
"content": [{
"type": "text",
"text": "Barber shaves all those, who do not shave themselves. Who shaves the barber?"
}
]
}, {
"role": "assistant",
"content": [{
"type": "text",
"text": "This is the barber paradox. Such a barber cannot logically exist."
}
],
"reasoning": "If the barber shaves himself, he should not; if he does not, he should.",
"reasoning_details": [{
"type": "reasoning.encrypted",
"data": "[REDACTED]"
}, {
"type": "reasoning.summary",
"summary": "Barber shaving himself creates contradiction"
}, {
"type": "reasoning.text",
"text": "If the barber shaves himself, he should not; if he does not, he should.",
"signature": "sig_123"
}
]
}, {
"role": "user",
"content": [{
"type": "text",
"text": "What if there are 2 barbers?"
}
]
}
],
"reasoning": {
"enabled": true,
"summary": "detailed",
"exclude": false
}
}
{
"id": "chatcmpl-Ae0TWsy2VPnSfBbv5UztnSdYUMFP3",
"choices": [
{
"message": {
"content": "Elastic is a company that provides a range of software solutions for search, logging, security, and analytics, built around the open-source Elasticsearch engine.",
"role": "assistant"
},
"finish_reason": "stop",
"index": 0
}
],
"model": "gpt-4o-2024-08-06",
"object": "chat.completion",
"usage": {
"completion_tokens": 34,
"prompt_tokens": 14,
"total_tokens": 48,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 0
}
}
}
{
"id": "chatcmpl-Ae0TWsy2VPnSfBbv5UztnSdYUMFP3",
"choices": [
{
"message": {
"role": "assistant",
"tool_calls": [
{
"index": 0,
"id": "call_KcAjWtAww20AihPHphUh46Gd",
"function": {
"arguments": "{\"item\":\"scarf\"}",
"name": "get_current_price"
},
"type": "function"
}
]
},
"finish_reason": "tool_calls",
"index": 0
}
],
"model": "gpt-4o-2024-08-06",
"object": "chat.completion",
"usage": {
"completion_tokens": 19,
"prompt_tokens": 135,
"total_tokens": 154,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 0
}
}
}
{
"id": "chatcmpl-910TWsy2VPnSfBbv5UztnSdYUJA10",
"choices": [
{
"message": {
"content": "With two barbers, the paradox disappears — each barber can shave the other.",
"role": "assistant",
"reasoning": "The contradiction only arises when a single barber must decide whether to shave himself. Two barbers eliminate the self-reference.",
"reasoning_details": [
{
"type": "reasoning.encrypted",
"data": "[REDACTED]"
},
{
"type": "reasoning.summary",
"summary": "Two barbers can shave each other, avoiding the self-reference paradox."
},
{
"type": "reasoning.text",
"text": "The contradiction only arises when a single barber must decide whether to shave himself. Two barbers eliminate the self-reference.",
"signature": "sig_example"
}
]
},
"finish_reason": "stop",
"index": 0
}
],
"model": "openai-gpt-oss-120b",
"object": "chat.completion",
"usage": {
"completion_tokens": 38,
"prompt_tokens": 74,
"total_tokens": 112,
"completion_tokens_details": {
"reasoning_tokens": 10
}
}
}