POST /responses accepts requests in OpenAI’s Responses API format and returns typed output items such as reasoning, message, and function_call. It works with every Venice text model and accepts an API key or x402 wallet auth.
When to use it
Use/responses when your client already speaks the Responses format, for example coding agents and SDKs built on OpenAI’s responses.create. It is the best way to run Codex-style agents on OpenAI models through Venice.
For non-OpenAI models, /chat/completions is still the most complete option: it supports structured outputs, file inputs, E2EE models, and every venice_parameters option on every model.
Quick start
output array of typed items and a usage object:
How requests are served
Venice serves each request in one of two modes.- Native. Requests for OpenAI models are forwarded to OpenAI’s Responses API without translation. This covers every
openai-*model exceptopenai-gpt-oss-120b. Response items, IDs, reasoning, and stream events come back exactly as OpenAI returns them. - Translated. Every other model, plus
openai-gpt-oss-120b, goes through a translation layer: Venice converts the request into a Chat Completions request, runs it, and converts the result back. Anything the Chat Completions format cannot express is ignored or rejected. See Limitations.
web_search tool, web_search: true, or venice_parameters.enable_web_search), x_search, venice_parameters.character_slug, and venice_parameters.enable_web_scraping. Requests with fields native mode cannot serve, listed under Native mode, fall back the same way.
You can control and inspect the mode with headers:
Conversations are stateless
Venice does not store responses, in either mode. Send the whole conversation ininput on every request, appending the previous output items and any tool results.
store is always treated as false. previous_response_id and conversation cannot be resolved because nothing is stored. With the default auto mode, such requests are served in translated mode, where those fields are ignored. With x-venice-responses-mode: native they return 400.
Native mode
Native mode supports the Responses features that OpenAI’s own endpoint does, including:instructions, applied as written.- Structured outputs with
text.format. OpenAI validates schemas strictly, so an invalid schema returns 400. For example,strict: truerequiresadditionalProperties: false, andjson_objectrequires the word “json” somewhere in the input. - Function tools and
tool_choicein OpenAI’s Responses shape{"type": "function", "name": "..."}, plus parallel tool calls. - Client-executed tools:
custom(freeform input),namespace,tool_search,apply_patch,shellandlocal_shellwith a local environment, and computer use. - Reasoning summaries, on by default. Reasoning models always return
encrypted_contenton reasoning items. Send those items back ininputon the next turn to keep the model’s reasoning context. - Prompt caching with
prompt_cache_key. Venice scopes the key to your account, so it never shares a cache with other users, and returns your original key in the response. - Image inputs, including
detail: "original", andinput_filewith inlinefile_data. - Pro models (
*-pro), which run OpenAI’s Pro reasoning mode automatically.
instructions apply unchanged. Set venice_parameters.include_venice_system_prompt: true to add it.
These need server-side storage or provider-hosted resources, so they are not available natively:
These restrictions also apply to tools loaded later in the conversation through
additional_tools or tool_search_output items. A Venice search tool declared there also routes the request to translated mode.
Unrecognized request fields are treated the same way: with auto the request is served in translated mode, and with native it returns 400.
Streaming
Setstream: true to receive server-sent events. Both modes end the stream with data: [DONE].
In native mode, events are exactly OpenAI’s, including response.reasoning_summary_text.delta for reasoning summaries and response.function_call_arguments.delta for tool arguments.
In translated mode, the events are response.created, response.output_item.added, response.content_part.added, response.output_text.delta, response.function_call_arguments.delta, response.content_part.done, response.output_item.done, and finally response.completed, response.incomplete, or response.failed. Two events differ from OpenAI’s:
- Reasoning text streams as
response.reasoning.delta, not as OpenAI’s reasoning summary events. - When web search runs, a
response.web_search.doneevent carries the search results before the answer streams.
Translated mode parameters
Web search is billed as search augmentation, as on
/chat/completions.
A few translated-mode behaviors to plan for:
instructionsare applied as a system message.- Custom tools return
custom_tool_callitems with the rawinput. If a model returns malformed custom-tool arguments, or calls a tool that is not declared in the request, the call comes back as an ordinaryfunction_callwith the original arguments, so your client can report a tool error and continue. - While a custom tool’s input streams, its call item and later items can arrive only after the input is complete. The stream sends SSE keep-alive comments while it waits.
- Image
detail: "original"is sent ashighto providers that do not support it.
Limitations
These gaps apply to translated mode: every non-OpenAI model,openai-gpt-oss-120b, and OpenAI model requests that use Venice-only features. Unless noted, the request still succeeds and the field is ignored.