Skip to main content
POST /responses accepts requests in OpenAI’s Responses API format and returns typed output items such as reasoning, message, and function_call. It works with every Venice text model and accepts an API key or x402 wallet auth.
This endpoint is in beta and available to all API users. OpenAI models are served natively, so they behave like OpenAI’s own Responses API. Every other model goes through a translation layer with some gaps. Read How requests are served and Limitations before building on it.

When to use it

Use /responses when your client already speaks the Responses format, for example coding agents and SDKs built on OpenAI’s responses.create. It is the best way to run Codex-style agents on OpenAI models through Venice. For non-OpenAI models, /chat/completions is still the most complete option: it supports structured outputs, file inputs, E2EE models, and every venice_parameters option on every model.

Quick start

The response contains an output array of typed items and a usage object:

How requests are served

Venice serves each request in one of two modes.
  • Native. Requests for OpenAI models are forwarded to OpenAI’s Responses API without translation. This covers every openai-* model except openai-gpt-oss-120b. Response items, IDs, reasoning, and stream events come back exactly as OpenAI returns them.
  • Translated. Every other model, plus openai-gpt-oss-120b, goes through a translation layer: Venice converts the request into a Chat Completions request, runs it, and converts the result back. Anything the Chat Completions format cannot express is ignored or rejected. See Limitations.
OpenAI model requests that use a Venice-only feature are served in translated mode so the feature keeps working. That includes Venice web search (a web_search tool, web_search: true, or venice_parameters.enable_web_search), x_search, venice_parameters.character_slug, and venice_parameters.enable_web_scraping. Requests with fields native mode cannot serve, listed under Native mode, fall back the same way. You can control and inspect the mode with headers:

Conversations are stateless

Venice does not store responses, in either mode. Send the whole conversation in input on every request, appending the previous output items and any tool results.
store is always treated as false. previous_response_id and conversation cannot be resolved because nothing is stored. With the default auto mode, such requests are served in translated mode, where those fields are ignored. With x-venice-responses-mode: native they return 400.

Native mode

Native mode supports the Responses features that OpenAI’s own endpoint does, including:
  • instructions, applied as written.
  • Structured outputs with text.format. OpenAI validates schemas strictly, so an invalid schema returns 400. For example, strict: true requires additionalProperties: false, and json_object requires the word “json” somewhere in the input.
  • Function tools and tool_choice in OpenAI’s Responses shape {"type": "function", "name": "..."}, plus parallel tool calls.
  • Client-executed tools: custom (freeform input), namespace, tool_search, apply_patch, shell and local_shell with a local environment, and computer use.
  • Reasoning summaries, on by default. Reasoning models always return encrypted_content on reasoning items. Send those items back in input on the next turn to keep the model’s reasoning context.
  • Prompt caching with prompt_cache_key. Venice scopes the key to your account, so it never shares a cache with other users, and returns your original key in the response.
  • Image inputs, including detail: "original", and input_file with inline file_data.
  • Pro models (*-pro), which run OpenAI’s Pro reasoning mode automatically.
The Venice system prompt is off by default in native mode, so your instructions apply unchanged. Set venice_parameters.include_venice_system_prompt: true to add it. These need server-side storage or provider-hosted resources, so they are not available natively: These restrictions also apply to tools loaded later in the conversation through additional_tools or tool_search_output items. A Venice search tool declared there also routes the request to translated mode. Unrecognized request fields are treated the same way: with auto the request is served in translated mode, and with native it returns 400.

Streaming

Set stream: true to receive server-sent events. Both modes end the stream with data: [DONE]. In native mode, events are exactly OpenAI’s, including response.reasoning_summary_text.delta for reasoning summaries and response.function_call_arguments.delta for tool arguments. In translated mode, the events are response.created, response.output_item.added, response.content_part.added, response.output_text.delta, response.function_call_arguments.delta, response.content_part.done, response.output_item.done, and finally response.completed, response.incomplete, or response.failed. Two events differ from OpenAI’s:
  • Reasoning text streams as response.reasoning.delta, not as OpenAI’s reasoning summary events.
  • When web search runs, a response.web_search.done event carries the search results before the answer streams.

Translated mode parameters

Web search is billed as search augmentation, as on /chat/completions. A few translated-mode behaviors to plan for:
  • instructions are applied as a system message.
  • Custom tools return custom_tool_call items with the raw input. If a model returns malformed custom-tool arguments, or calls a tool that is not declared in the request, the call comes back as an ordinary function_call with the original arguments, so your client can report a tool error and continue.
  • While a custom tool’s input streams, its call item and later items can arrive only after the input is complete. The stream sends SSE keep-alive comments while it waits.
  • Image detail: "original" is sent as high to providers that do not support it.

Limitations

These gaps apply to translated mode: every non-OpenAI model, openai-gpt-oss-120b, and OpenAI model requests that use Venice-only features. Unless noted, the request still succeeds and the field is ignored.