> ## Documentation Index
> Fetch the complete documentation index at: https://docs.cowagent.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI

> OpenAI model configuration (Text / Vision / Image / Speech / Embedding)

OpenAI offers the most complete coverage and can simultaneously serve text chat, vision understanding, image generation, speech-to-text (ASR), text-to-speech (TTS), and embedding. A single `open_ai_api_key` lets the Agent use all of these capabilities.

<Tip>
  All capabilities below can be configured in one place via the "Model Management" page in the Web Console, with no need to manually edit the configuration file.
</Tip>

## Text Chat

```json theme={null}
{
  "model": "gpt-6.1-sol",
  "open_ai_api_key": "YOUR_API_KEY",
  "open_ai_api_base": "https://api.openai.com/v1"
}
```

| Parameter | Description |
| - | - |
| `model` | Same as OpenAI's [model parameter](https://platform.openai.com/docs/models); supports `gpt-6.1-sol`, `gpt-6-luna`, `gpt-6-sol`, `gpt-6-astra`, `gpt-5.6-luna`, `gpt-5.6-terra`, `gpt-5.6-sol`, `gpt-5.5`, `gpt-5.4`, `gpt-5.4-mini`, `gpt-5.4-nano`, the `gpt-5` series, `gpt-4.1`, etc. Agent mode defaults to `gpt-6.1-sol`; `gpt-6-astra` is the top flagship (most intelligent, higher cost); tool calling for the `gpt-6` / `gpt-6.1` series goes through the Responses API; use `gpt-5.4` for better cost-efficiency |
| `open_ai_api_key` | Create one on the [OpenAI Platform](https://platform.openai.com/api-keys) |
| `open_ai_api_base` | Optional; change it to access a third-party proxy |
| `open_ai_api_type` | Optional API protocol, defaults to `auto` (only the `gpt-6` / `gpt-6.1` series uses the Responses API, others use Chat Completions). If your endpoint no longer serves `/chat/completions`, set it to `responses` to send all requests to `/responses`; `chat` always uses `/chat/completions`. For Docker deployments, set it via the `OPEN_AI_API_TYPE` environment variable |
| `bot_type` | Not required when using OpenAI's official models; set to `openai` when accessing other providers via the compatible protocol |

## Image Understanding

OpenAI models like `gpt-5.5`, `gpt-5.4`, `gpt-4o`, and `gpt-4.1` natively support vision. Once `open_ai_api_key` is configured, the Agent's Vision tool automatically uses the main model to recognize images. If the main model does not support vision or you want to specify it explicitly, set it in the configuration file:

```json theme={null}
{
  "tools": {
    "vision": {
      "model": "gpt-5.4-mini"
    }
  }
}
```

Supported Vision models: `gpt-5.5`, `gpt-5.4`, `gpt-5.4-mini`, `gpt-5.4-nano`, `gpt-5`, `gpt-4.1`, `gpt-4.1-mini`, `gpt-4o`.

## Image Generation

Specify the image generation model in the configuration file; the Agent automatically routes image generation skill calls to OpenAI:

```json theme={null}
{
  "skills": {
    "image-generation": {
      "model": "gpt-image-2.5-flare"
    }
  }
}
```

Supported image generation models: `gpt-image-2.5-flare` (recommended default), `gpt-image-2.5-sunburst`, `gpt-image-2`, `gpt-image-1`.

## Speech-to-Text (ASR)

```json theme={null}
{
  "voice_to_text": "openai",
  "voice_to_text_model": "gpt-4o-mini-transcribe"
}
```

| Parameter | Description |
| - | - |
| `voice_to_text` | Set to `openai` to enable OpenAI speech-to-text |
| `voice_to_text_model` | Optional, defaults to `gpt-4o-mini-transcribe`; can also be `gpt-4o-transcribe`, `whisper-1` |

Credentials are automatically reused from `open_ai_api_key`.

## Text-to-Speech (TTS)

```json theme={null}
{
  "text_to_voice": "openai",
  "text_to_voice_model": "tts-1",
  "tts_voice_id": "alloy"
}
```

| Parameter | Description |
| - | - |
| `text_to_voice_model` | `tts-1`, `tts-1-hd`, `gpt-4o-mini-tts` |
| `tts_voice_id` | Voices: `alloy`, `echo`, `fable`, `onyx`, `nova`, `shimmer`, `ash`, `ballad`, `coral`, `sage`, `verse` |

## Embedding

```json theme={null}
{
  "embedding_provider": "openai",
  "embedding_model": "text-embedding-3-small"
}
```

Available models: `text-embedding-3-small`, `text-embedding-3-large`, `text-embedding-ada-002`. After changing the embedding, run `/memory rebuild-index` to rebuild the index.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.