Skip to content
agentgateway has joined the Agentic AI Foundation — Learn more

For the complete documentation index, see llms.txt. Markdown versions of all docs pages are available by appending .md to any docs URL.

OpenAI-compatible providers

Verified Code examples on this page have been automatically tested and verified.
Page as Markdown

Configure OpenAI-compatible providers with AgentgatewayModel presets or AgentgatewayBackend resources.

Configure OpenAI-compatible LLM providers with an AgentgatewayModel or AgentgatewayBackend resource.

Overview

Choose the resource that matches how you want to route requests. Both approaches use the provider credentials that you create on this page.

ResourceWhen to use it
AgentgatewayModelRoute by the model name in the request and serve standard LLM paths, such as /v1/chat/completions. Provider presets supply the default URL and supported request formats.
AgentgatewayBackend with an HTTPRouteConfigure routing by path, header, or other HTTPRoute matches. For the providers on this page, set spec.ai.provider.openai and configure the upstream host and path.

The AgentgatewayModel API is enabled by default in agentgateway 1.6. For more information about the two approaches, see About models.

Built-in OpenAI-compatible providers

For an AgentgatewayModel, use the case-sensitive spec.provider value in the table. You can omit spec.baseURL to use the preset’s default URL. For an AgentgatewayBackend, use the host and path columns with port: 443.

ProviderModel spec.providerBackend hostBackend path
BasetenBaseteninference.baseten.co/v1/chat/completions
CerebrasCerebrasapi.cerebras.ai/v1/chat/completions
CohereCohereapi.cohere.ai/compatibility/v1/chat/completions
DeepInfraDeepinfraapi.deepinfra.com/v1/openai/chat/completions
DeepSeekDeepseekapi.deepseek.com/v1/chat/completions
Fireworks AIFireworksapi.fireworks.ai/inference/v1/chat/completions
GroqGroqapi.groq.com/openai/v1/chat/completions
Hugging FaceHuggingfacerouter.huggingface.co/v1/chat/completions
MetaMetaapi.meta.ai/v1/chat/completions
Mistral AIMistralapi.mistral.ai/v1/chat/completions
OpenRouterOpenrouteropenrouter.ai/api/v1/chat/completions
Together AITogetherAIapi.together.xyz/v1/chat/completions
xAIXAIapi.x.ai/v1/chat/completions

If your provider is not in this list but still exposes the OpenAI Chat Completions API, use the generic endpoint template. If the upstream does not match the OpenAI API format, use custom providers instead.

Before you begin

Install and set up an agentgateway proxy.

Set up provider credentials

Create a Secret with an API key for the provider that you want to use. The model and backend examples both reference this Secret. Complete one configuration path after you create the credentials.

  1. Get an API key for your provider. For Meta, create a key in the Meta Model API dashboard.

  2. Save the API key in an environment variable.

    export MY_API_KEY='<your-api-key>'
  3. Create a Kubernetes secret to store your API key.

    kubectl apply -f- <<EOF
    apiVersion: v1
    kind: Secret
    metadata:
      name: llm-provider-secret
      namespace: agentgateway-system
    type: Opaque
    stringData:
      Authorization: $MY_API_KEY
    EOF

Configure an AgentgatewayModel

Use a provider preset to serve standard LLM APIs through the gateway. Each example uses the provider’s default URL and the Secret from the previous section. The model names in the tabs are examples. Use a model that your provider supports.

  1. Enable LLM serving on a listener. The listener must allow the AgentgatewayModel route kind. The examples below attach to the http listener of the agentgateway-proxy Gateway.

  2. Create the model. Select the tab for your provider and use that provider’s API key in the Secret. Set spec.match.model to the model that you want to serve.

    kubectl apply -f- <<EOF
    apiVersion: agentgateway.dev/v1alpha1
    kind: AgentgatewayModel
    metadata:
      name: llm-model
      namespace: agentgateway-system
    spec:
      parentRefs:
      - group: gateway.networking.k8s.io
        kind: Gateway
        name: agentgateway-proxy
        sectionName: http
      match:
        model: meta-llama/Llama-3.1-8B-Instruct
      provider: Baseten
      policies:
        auth:
          secretRef:
            name: llm-provider-secret
    EOF
    SettingDescription
    spec.parentRefsAttaches the model to the http listener of agentgateway-proxy in the same namespace. This listener serves the model without a separate backend or HTTPRoute.
    spec.match.modelThe model name that clients send in requests. Each example forwards the same name to the provider.
    spec.providerThe provider preset, such as Meta or Groq. The preset supplies the URL and request formats.
    spec.policies.auth.secretRef.nameThe Secret that contains the provider API key in its Authorization entry. The Secret must be in the model’s namespace.

    To override a preset’s URL, set spec.baseURL. For the URL format and examples, see Providers.

  3. Send a request through the gateway. The model endpoint is /v1/chat/completions. Replace <your-model> with the spec.match.model value from your model configuration.

    curl "http://$INGRESS_GW_ADDRESS/v1/chat/completions" \
      -H 'Content-Type: application/json' \
      -d '{
        "model": "<your-model>",
        "messages": [{"role": "user", "content": "Explain retrieval-augmented generation in one sentence."}]
      }' | jq

Configure an AgentgatewayBackend

Use this approach to select the backend with an HTTPRoute. The examples configure Chat Completions through spec.ai.provider.openai and expose the backend on /llm. The model names in the tabs are examples. Use a model that your provider supports.

  1. Create an AgentgatewayBackend resource that points the openai provider at your provider’s host and path. Select the tab for your provider.

    kubectl apply -f- <<EOF
    apiVersion: agentgateway.dev/v1alpha1
    kind: AgentgatewayBackend
    metadata:
      name: llm-backend
      namespace: agentgateway-system
    spec:
      ai:
        provider:
          openai:
            model: meta-llama/Llama-3.1-8B-Instruct
          host: inference.baseten.co
          port: 443
          path: /v1/chat/completions
      policies:
        auth:
          secretRef:
            name: llm-provider-secret
        tls:
          sni: inference.baseten.co
    EOF

    Review the following table to understand this configuration.

    SettingDescription
    ai.provider.openai.modelOptional upstream model override. Omit this parameter to pass the client-provided model through.
    host and portThe provider’s API host and port. Use 443 for HTTPS endpoints.
    pathThe provider’s chat completions path. Omit this parameter for providers that use the standard /v1/chat/completions path.
    policies.auth.secretRefReferences the secret that contains your provider API key.
    policies.tls.sniEnables TLS and sets the SNI value to the upstream hostname.
  2. Create an HTTPRoute resource that routes incoming traffic to the AgentgatewayBackend.

    kubectl apply -f- <<EOF
    apiVersion: gateway.networking.k8s.io/v1
    kind: HTTPRoute
    metadata:
      name: llm-route
      namespace: agentgateway-system
    spec:
      parentRefs:
        - name: agentgateway-proxy
          namespace: agentgateway-system
      rules:
      - matches:
        - path:
            type: PathPrefix
            value: /llm
        backendRefs:
        - name: llm-backend
          namespace: agentgateway-system
          group: agentgateway.dev
          kind: AgentgatewayBackend
    EOF
  3. Send a request to verify the setup. Replace the model value with the model that you configured on the AgentgatewayBackend.

    curl "$INGRESS_GW_ADDRESS/llm" -H content-type:application/json -d '{
       "model": "<your-model>",
       "messages": [
         {
           "role": "user",
           "content": "Explain retrieval-augmented generation in one sentence."
         }
       ]
     }' | jq

Other OpenAI-compatible providers

Perplexity example

Perplexity exposes an OpenAI-compatible API for search-augmented models and uses the standard chat completions path, so you do not need to set path.

kubectl apply -f- <<EOF
apiVersion: agentgateway.dev/v1alpha1
kind: AgentgatewayBackend
metadata:
  name: perplexity
  namespace: agentgateway-system
spec:
  ai:
    provider:
      openai:
        model: sonar
      host: api.perplexity.ai
      port: 443
  policies:
    auth:
      secretRef:
        name: perplexity-secret
    tls:
      sni: api.perplexity.ai
EOF

Generic OpenAI-compatible endpoint

Use this template when the provider exposes the OpenAI Chat Completions API but is not listed in the Built-in OpenAI-compatible providers table.

apiVersion: agentgateway.dev/v1alpha1
kind: AgentgatewayBackend
metadata:
  name: generic-openai
  namespace: agentgateway-system
spec:
  ai:
    provider:
      openai:
        model: <upstream-model-name>
      host: api.example.com
      port: 443
      path: /v1/chat/completions
  policies:
    auth:
      secretRef:
        name: provider-secret
    tls:
      sni: api.example.com

Use the following fields to adapt the template:

SettingDescription
ai.provider.openai.modelOptional upstream model override. Omit this parameter to pass the client-provided model through.
host and portRequired target address for the external provider endpoint.
pathThe provider’s chat completions path. Omit this parameter for the standard /v1/chat/completions path.
policies.authAttach the provider API key secret to outbound requests.
policies.tls.sniEnable TLS and set the SNI value to the upstream hostname.

If the upstream needs mixed API formats or a cluster-local backend target, use custom providers instead. For self-hosted targets that already have guides, prefer the dedicated Ollama and vLLM pages.

Next steps

Was this page helpful?
Agentgateway assistant

Ask me anything about agentgateway configuration, features, or usage.

Note: AI-generated content might contain errors; please verify and test all returned information.

Tip: one topic per conversation gives the best results. Use the + button in the chat header to start a new conversation.

Switching topics? Starting a new conversation improves accuracy.
↑↓ navigate ↵ select esc dismiss

What could be improved?

Your feedback helps us improve assistant answers and identify docs gaps we should fix.

Need more help? Join us on Discord: https://discord.gg/y9efgEmppm

Want to use your own agent? Add the Solo MCP server to query our docs directly. Get started here: https://search.solo.io/.