Perform non-streaming chat completion inference on the service Generally available

POST /_inference/chat_completion/{inference_id}

The chat completion inference API enables rich responses for chat completion tasks. It only works with the chat_completion task type.

NOTE: The chat_completion task type supports both streaming and non-streaming. The Chat completion inference API provides more comprehensive customization options through more fields and function calling support. To determine whether a given inference service supports this task type, please see the page for that service.

These services support non-streaming chat completion inference:

  • AI21
  • Azure OpenAI
  • Deepseek
  • Elastic
  • FireworksAI
  • Groq
  • Huggingface
  • IBMWatsonX
  • Llama
  • Mistral
  • NVIDIA
  • OpenAI
  • OpenShiftAI

Path parameters

  • inference_id string Required

    The inference Id

Query parameters

  • timeout string

    Specifies the amount of time to wait for the inference request to complete.

    Default value is 120s.

application/json

Body Required

  • messages array[object] Required

    A list of objects representing the conversation. Requests should generally only add new messages from the user (role user). The other message roles (assistant, system, or tool) should generally only be copied from the response to a previous completion request, such that the messages array is built up throughout a conversation.

    Hide messages attributes Show messages attributes object

    An object representing part of the conversation.

    • content string | array[object]

      The content of the message.

      String example:

      {
         "content": "Some string"
      }
      

      Text example:

      {
        "content": [
            {
             "text": "Some text",
             "type": "text"
            }
         ]
      }
      

      Image example:

      {
        "content": [
            {
             "image_url": {
               "url": "data:image/jpeg;base64,..."
             },
             "type": "image_url"
            }
         ]
      }
      

      File example:

      {
        "content": [
            {
             "file": {
               "file_data": "data:application/pdf;base64,...",
               "filename": "somePDF"
             },
             "type": "file"
            }
         ]
      }
      
      One of:

      An object style representation of a single portion of a conversation.

    • role string Required

      The role of the message author. Valid values are user, assistant, system, and tool.

    • tool_call_id string

      Only for tool role messages. The tool call that this message is responding to.

    • tool_calls array[object]

      Only for assistant role messages. The tool calls generated by the model. If it's specified, the content field is optional. Example:

      {
        "tool_calls": [
            {
                "id": "call_KcAjWtAww20AihPHphUh46Gd",
                "type": "function",
                "function": {
                    "name": "get_current_weather",
                    "arguments": "{\"location\":\"Boston, MA\"}"
                }
            }
        ]
      }
      
      Hide tool_calls attributes Show tool_calls attributes object

      A tool call generated by the model.

      • id string Required

        The identifier of the tool call.

      • function object Required

        The function that the model called.

        Hide function attributes Show function attributes object
        • arguments string Required

          The arguments to call the function with in JSON format.

        • name string Required

          The name of the function to call.

      • type string Required

        The type of the tool call.

    • reasoning string

      Only for assistant role messages. The reasoning details generated by the model as plaintext. Currently supported only for elastic provider.

    • reasoning_details array[object]

      Only for assistant role messages. The reasoning details generated by the model as structured data. Currently supported only for elastic provider.

      Type representing the different types of reasoning details that can be included in the response from the model. Currently supported only for elastic provider.

      One of:
  • model string

    The ID of the model to use. By default, the model ID is set to the value included when creating the inference endpoint.

  • max_completion_tokens number

    The upper bound limit for the number of tokens that can be generated for a completion request.

  • reasoning object

    The reasoning configuration for the completion request. This controls the model's reasoning process in one of two ways:

    • By specifying the model’s reasoning effort level with the effort field.
    • By enabling reasoning with default settings by setting enabled field to true.

    It also includes optional settings to control:

    • The level of detail in the summary returned in the response with the summary field.
    • Whether reasoning details are included in the response at all with the exclude field.

    Example (effort):

    {
       "reasoning": {
           "effort": "high",
           "summary": "concise",
           "exclude": false
       }
    }
    

    Example (enabled):

    {
       "reasoning": {
           "enabled": true,
           "summary": "concise",
           "exclude": false
       }
    }
    

    Currently supported only for elastic provider.

    Hide reasoning attributes Show reasoning attributes object
    • effort string

      The level of effort the model should put into reasoning. This is a hint that guides the model in how much effort to put into reasoning, with xhigh being the most effort and none being no effort.

      Values are xhigh, high, medium, low, minimal, or none.

    • enabled boolean

      Whether to enable reasoning with default settings. This is a shortcut for enabling reasoning without having to specify the other parameters. If enabled is set to true, then reasoning at the medium effort level is enabled. Ignored if effort is specified, in which case that parameter will control the reasoning process instead.

    • exclude boolean

      Whether to exclude reasoning information from the response. If true, the response will not include any reasoning details.

    • summary string

      The level of detail included in the reasoning summary returned in the response. This is a hint on how much detail to include in the summary of the reasoning that is returned in the response, with auto being the default level of detail, concise being less detail, and detailed being more detail.

      Values are auto, concise, or detailed.

  • stop array[string]

    A sequence of strings to control when the model should stop generating additional tokens.

  • temperature number

    The sampling temperature to use.

  • tool_choice string | object

    Controls which tool is called by the model. String representation: One of auto, none, or requrired. auto allows the model to choose between calling tools and generating a message. none causes the model to not call any tools. required forces the model to call one or more tools. Example (object representation):

    {
      "tool_choice": {
          "type": "function",
          "function": {
              "name": "get_current_weather"
          }
      }
    }
    
    One of:
  • tools array[object]

    A list of tools that the model can call. Example:

    {
      "tools": [
          {
              "type": "function",
              "function": {
                  "name": "get_price_of_item",
                  "description": "Get the current price of an item",
                  "parameters": {
                      "type": "object",
                      "properties": {
                          "item": {
                              "id": "12345"
                          },
                          "unit": {
                              "type": "currency"
                          }
                      }
                  }
              }
          }
      ]
    }
    
    Hide tools attributes Show tools attributes object

    A list of tools that the model can call.

    • type string Required

      The type of tool.

    • function object Required

      The function definition.

      Hide function attributes Show function attributes object
      • description string

        A description of what the function does. This is used by the model to choose when and how to call the function.

      • name string Required

        The name of the function.

      • parameters object

        The parameters the functional accepts. This should be formatted as a JSON object.

      • strict boolean

        Whether to enable schema adherence when generating the function call.

  • top_p number

    Nucleus sampling, an alternative to sampling with temperature.

Responses

  • 200 application/json
    Hide response attributes Show response attributes object
    • id string Required

      The unique identifier for the completion.

    • choices array[object]

      The list of completion choices the model generated for the input message.

      Hide choices attributes Show choices attributes object

      A single completion choice returned by the model.

      • message object Required

        The message generated by the model for this choice.

        Hide message attributes Show message attributes object
        • content string

          The content of the message.

        • refusal string

          The refusal message generated by the model.

        • role string

          The role of the message author.

        • reasoning string

          The reasoning generated by the model as plaintext. Currently supported only for the elastic provider.

        • tool_calls array[object]

          The tool calls generated by the model.

          A tool call generated by the model.

          A tool call generated by the model.

        • reasoning_details array[object]

          The reasoning details generated by the model as structured data. Currently supported only for the elastic provider.

      • finish_reason string

        The reason the model stopped generating tokens. Common values are stop (natural stopping point) and tool_calls (the model called a tool). Omitted when the reason is not available.

      • index number Required

        The index of this choice in the list of choices.

    • model string Required

      The model used to generate the completion.

    • object string Required

      The object type.

    • usage object

      The token usage statistics for the completion request. Omitted when usage information is not available.

      Hide usage attributes Show usage attributes object
      • completion_tokens number Required

        The number of tokens in the generated completion.

      • prompt_tokens number Required

        The number of tokens in the prompt.

      • total_tokens number Required

        The total number of tokens used (prompt + completion).

      • prompt_tokens_details object

        Breakdown of the tokens used in the prompt. Omitted when no details are available.

        Hide prompt_tokens_details attributes Show prompt_tokens_details attributes object
        • cached_tokens number

          The number of tokens that were cached from a previous request.

        • cache_write_tokens number

          The number of tokens written to the cache.

      • completion_tokens_details object

        Breakdown of the tokens used in the completion. Omitted when no details are available.

        Hide completion_tokens_details attribute Show completion_tokens_details attribute object
        • reasoning_tokens number

          The number of tokens used for reasoning by the model.

POST /_inference/chat_completion/{inference_id}
POST _inference/chat_completion/openai-completion
{
  "model": "gpt-4o",
  "messages": [
      {
          "role": "user",
          "content": "What is Elastic?"
      }
  ]
}
resp = client.inference.non_streaming_chat_completion(
    inference_id="openai-completion",
    chat_completion_request={
        "model": "gpt-4o",
        "messages": [
            {
                "role": "user",
                "content": "What is Elastic?"
            }
        ]
    },
)
const response = await client.inference.nonStreamingChatCompletion({
  inference_id: "openai-completion",
  chat_completion_request: {
    model: "gpt-4o",
    messages: [
      {
        role: "user",
        content: "What is Elastic?",
      },
    ],
  },
});
response = client.inference.non_streaming_chat_completion(
  inference_id: "openai-completion",
  body: {
    "model": "gpt-4o",
    "messages": [
      {
        "role": "user",
        "content": "What is Elastic?"
      }
    ]
  }
)
$resp = $client->inference()->nonStreamingChatCompletion([
    "inference_id" => "openai-completion",
    "body" => [
        "model" => "gpt-4o",
        "messages" => array(
            [
                "role" => "user",
                "content" => "What is Elastic?",
            ],
        ),
    ],
]);
curl -X POST -H "Authorization: ApiKey $ELASTIC_API_KEY" -H "Content-Type: application/json" -d '{"model":"gpt-4o","messages":[{"role":"user","content":"What is Elastic?"}]}' "$ELASTICSEARCH_URL/_inference/chat_completion/openai-completion"
Request examples
Run `POST _inference/chat_completion/openai-completion` to perform a chat completion on the example question with non-streaming.
{
  "model": "gpt-4o",
  "messages": [
      {
          "role": "user",
          "content": "What is Elastic?"
      }
  ]
}
Run `POST _inference/chat_completion/openai-completion` to perform a non-streaming chat completion using an Assistant message with `tool_calls`.
{
  "messages": [
      {
          "role": "assistant",
          "content": "Let's find out what the weather is",
          "tool_calls": [ 
              {
                  "id": "call_KcAjWtAww20AihPHphUh46Gd",
                  "type": "function",
                  "function": {
                      "name": "get_current_weather",
                      "arguments": "{\"location\":\"Boston, MA\"}"
                  }
              }
          ]
      },
      { 
          "role": "tool",
          "content": "The weather is cold",
          "tool_call_id": "call_KcAjWtAww20AihPHphUh46Gd"
      }
  ]
}
Run `POST _inference/chat_completion/openai-completion` to perform a non-streaming chat completion using a User message with `tools` and `tool_choice`.
{
  "messages": [
      {
          "role": "user",
          "content": [
              {
                  "type": "text",
                  "text": "What's the price of a scarf?"
              }
          ]
      }
  ],
  "tools": [
      {
          "type": "function",
          "function": {
              "name": "get_current_price",
              "description": "Get the current price of a item",
              "parameters": {
                  "type": "object",
                  "properties": {
                      "item": {
                          "id": "123"
                      }
                  }
              }
          }
      }
  ],
  "tool_choice": {
      "type": "function",
      "function": {
          "name": "get_current_price"
      }
  }
}
Run `POST _inference/chat_completion/reasoning-chat-completion` to perform a non-streaming chat completion task, using both `effort` parameter based reasoning configuration and including reasoning generated by the model on previous step.
{
  "messages": [{
      "role": "user",
      "content": [{
          "type": "text",
          "text": "Barber shaves all those, who do not shave themselves. Who shaves the barber?"
        }
      ]
    }, {
      "role": "assistant",
      "content": [{
          "type": "text",
          "text": "This is the barber paradox. Such a barber cannot logically exist."
        }
      ],
      "reasoning": "If the barber shaves himself, he should not; if he does not, he should.",
      "reasoning_details": [{
          "type": "reasoning.encrypted",
          "data": "[REDACTED]"
        }, {
          "type": "reasoning.summary",
          "summary": "Barber shaving himself creates contradiction"
        }, {
          "type": "reasoning.text",
          "text": "If the barber shaves himself, he should not; if he does not, he should.",
          "signature": "sig_123"
        }
      ]
    }, {
      "role": "user",
      "content": [{
          "type": "text",
          "text": "What if there are 2 barbers?"
        }
      ]
    }
  ],
  "reasoning": {
    "effort": "high",
    "summary": "detailed",
    "exclude": false
  }
}
Run `POST _inference/chat_completion/reasoning-chat-completion` to perform a non-streaming chat completion task, using both `enabled` parameter based reasoning configuration and including reasoning generated by the model on previous step.
{
  "messages": [{
      "role": "user",
      "content": [{
          "type": "text",
          "text": "Barber shaves all those, who do not shave themselves. Who shaves the barber?"
        }
      ]
    }, {
      "role": "assistant",
      "content": [{
          "type": "text",
          "text": "This is the barber paradox. Such a barber cannot logically exist."
        }
      ],
      "reasoning": "If the barber shaves himself, he should not; if he does not, he should.",
      "reasoning_details": [{
          "type": "reasoning.encrypted",
          "data": "[REDACTED]"
        }, {
          "type": "reasoning.summary",
          "summary": "Barber shaving himself creates contradiction"
        }, {
          "type": "reasoning.text",
          "text": "If the barber shaves himself, he should not; if he does not, he should.",
          "signature": "sig_123"
        }
      ]
    }, {
      "role": "user",
      "content": [{
          "type": "text",
          "text": "What if there are 2 barbers?"
        }
      ]
    }
  ],
  "reasoning": {
    "enabled": true,
    "summary": "detailed",
    "exclude": false
  }
}
Response examples (200)
A successful response from `POST _inference/chat_completion/openai-completion` for a plain user message. The `object` field is `chat.completion` (as opposed to `chat.completion.chunk` in streaming responses). The `message` wrapper contains the model-generated text. Usage statistics include per-category token breakdowns when available.
{
  "id": "chatcmpl-Ae0TWsy2VPnSfBbv5UztnSdYUMFP3",
  "choices": [
    {
      "message": {
        "content": "Elastic is a company that provides a range of software solutions for search, logging, security, and analytics, built around the open-source Elasticsearch engine.",
        "role": "assistant"
      },
      "finish_reason": "stop",
      "index": 0
    }
  ],
  "model": "gpt-4o-2024-08-06",
  "object": "chat.completion",
  "usage": {
    "completion_tokens": 34,
    "prompt_tokens": 14,
    "total_tokens": 48,
    "prompt_tokens_details": {
      "cached_tokens": 0
    },
    "completion_tokens_details": {
      "reasoning_tokens": 0
    }
  }
}
A successful response from `POST _inference/chat_completion/openai-completion` when the model decides to call a tool. The `message` contains `tool_calls` instead of `content`, and `finish_reason` is `tool_calls`. Each tool call includes an `index`, an `id`, the function `name` and `arguments` (as a JSON string), and a `type`.
{
  "id": "chatcmpl-Ae0TWsy2VPnSfBbv5UztnSdYUMFP3",
  "choices": [
    {
      "message": {
        "role": "assistant",
        "tool_calls": [
          {
            "index": 0,
            "id": "call_KcAjWtAww20AihPHphUh46Gd",
            "function": {
              "arguments": "{\"item\":\"scarf\"}",
              "name": "get_current_price"
            },
            "type": "function"
          }
        ]
      },
      "finish_reason": "tool_calls",
      "index": 0
    }
  ],
  "model": "gpt-4o-2024-08-06",
  "object": "chat.completion",
  "usage": {
    "completion_tokens": 19,
    "prompt_tokens": 135,
    "total_tokens": 154,
    "prompt_tokens_details": {
      "cached_tokens": 0
    },
    "completion_tokens_details": {
      "reasoning_tokens": 0
    }
  }
}
A successful response from `POST _inference/chat_completion/reasoning-chat-completion` when the model returns structured reasoning data alongside its answer. The `message` includes a `reasoning` plaintext summary and a `reasoning_details` array covering all three detail types: `reasoning.encrypted` (opaque data), `reasoning.summary` (human-readable summary), and `reasoning.text` (full reasoning text with a verification signature). Currently supported only for the `elastic` provider.
{
  "id": "chatcmpl-910TWsy2VPnSfBbv5UztnSdYUJA10",
  "choices": [
    {
      "message": {
        "content": "With two barbers, the paradox disappears — each barber can shave the other.",
        "role": "assistant",
        "reasoning": "The contradiction only arises when a single barber must decide whether to shave himself. Two barbers eliminate the self-reference.",
        "reasoning_details": [
          {
            "type": "reasoning.encrypted",
            "data": "[REDACTED]"
          },
          {
            "type": "reasoning.summary",
            "summary": "Two barbers can shave each other, avoiding the self-reference paradox."
          },
          {
            "type": "reasoning.text",
            "text": "The contradiction only arises when a single barber must decide whether to shave himself. Two barbers eliminate the self-reference.",
            "signature": "sig_example"
          }
        ]
      },
      "finish_reason": "stop",
      "index": 0
    }
  ],
  "model": "openai-gpt-oss-120b",
  "object": "chat.completion",
  "usage": {
    "completion_tokens": 38,
    "prompt_tokens": 74,
    "total_tokens": 112,
    "completion_tokens_details": {
      "reasoning_tokens": 10
    }
  }
}