Skip to content
agentgateway has joined the Agentic AI Foundation — Learn more

For the complete documentation index, see llms.txt. Markdown versions of all docs pages are available by appending .md to any docs URL.

Responses

Page as Markdown

Send requests through agentgateway using the OpenAI Responses API.

The OpenAI Responses API (/v1/responses) is OpenAI’s interface for stateful, multi-step model interactions.

About

The OpenAI Responses API is a unified interface that supports text and multimodal generation, built-in tools, and multi-turn conversation state. Agentgateway proxies these requests to your configured providers while providing token usage tracking, observability metrics, and policy enforcement.

A provider that advertises the responses format also serves clients that send the Anthropic Messages format. Agentgateway converts a Messages request into a Responses request, and converts the buffered or streamed reply back. For the conversion order and its limits, see Provider format conversion.

Namespaced tools

Responses requests can include namespace tools. When the selected provider takes Chat Completions or Bedrock Converse requests, the conversion flattens each namespace member into a function name in the namespace__function format. The conversion restores buffered and streamed replies to the Responses shape. Function calls keep the original namespace and name fields.

A function tool choice can use the bare member name only when that name is unique across the namespace tools. When more than one namespace has a function with the same name, set the tool choice function name to namespace__function. Chat Completions and Bedrock conversions cannot enforce allowed_tools choices or convert custom namespace members. Requests that use those features fail before they reach the provider.

Route type configuration

In the simplified llm configuration, agentgateway automatically maps /v1/responses requests to the responses route type, so no explicit route configuration is required.

# yaml-language-server: $schema=https://agentgateway.dev/schema/config
llm:
  models:
  - name: "*"
    provider: openAI
    params:
      apiKey: "$OPENAI_API_KEY"

To configure the route type explicitly, use the gateways and routes format and set the responses route type in the policies.ai.routes map.

# yaml-language-server: $schema=https://agentgateway.dev/schema/config
gateways:
  default:
    port: 4000
routes:
- backends:
  - ai:
      name: openai
      provider:
        openAI: {}
  policies:
    ai:
      routes:
        "/v1/responses": "responses"
    backendAuth:
      key: "$OPENAI_API_KEY"

Note

For detailed information about model routing and configuration modes, see Model routing and aliases.

Usage in converted replies

Responses replies follow OpenAI usage conventions, even when agentgateway converts the upstream provider response from another format. When a provider reports prompt-cache tokens, usage.input_tokens includes those tokens. When they are available, cache counts are also reported separately in usage.input_tokens_details.cached_tokens and usage.input_tokens_details.cache_write_tokens.

Using the API

Using the Responses API works exactly the same as consuming OpenAI directly, with only a change to the base URL. This allows you to continue using existing code and SDKs.

Use HTTP POST for Responses requests through agentgateway, because the Responses WebSocket transport is not supported. With the simplified llm configuration, a WebSocket upgrade request to /v1/responses returns a 405 Method Not Allowed error with the websocket_not_supported code. With the routes format, when policies.ai.routes maps /v1/responses to responses, the upgrade request fails with a 400 error instead.

curl 'http://localhost:4000/v1/responses' \
--header 'Content-Type: application/json' \
--data '{
  "model": "gpt-4o-mini",
  "input": "Tell me a story"
}'
Was this page helpful?
Agentgateway assistant

Ask me anything about agentgateway configuration, features, or usage.

Note: AI-generated content might contain errors; please verify and test all returned information.

Tip: one topic per conversation gives the best results. Use the + button in the chat header to start a new conversation.

Switching topics? Starting a new conversation improves accuracy.
↑↓ navigate ↵ select esc dismiss

What could be improved?

Your feedback helps us improve assistant answers and identify docs gaps we should fix.

Need more help? Join us on Discord: https://discord.gg/y9efgEmppm

Want to use your own agent? Add the Solo MCP server to query our docs directly. Get started here: https://search.solo.io/.