> ## Documentation Index
> Fetch the complete documentation index at: https://docs.qredence.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Configure models

> Point Fleet RLM's Root and Sub model roles at any OpenAI-compatible Chat Completions endpoint.

Fleet RLM calls models through a stock `dspy.LM` against an OpenAI-compatible Chat Completions endpoint. You configure the Root and Sub roles separately in `config/fleet.toml`.

## Shipped configuration

The shipped configuration uses a DeepSeek model through the Databricks Unity AI Gateway for both roles.

```toml config/fleet.toml theme={null}
[llm.root]
model = "uscentral.ai_gateway.deepseek-v4-1-flash-service"
api_key_env = "DATABRICKS_TOKEN"
base_url_env = "FLEET_LLM_BASE_URL"
max_tokens = 16384
timeout_seconds = 300
num_retries = 3
cache = false

[llm.sub]
model = "uscentral.ai_gateway.deepseek-v4-1-flash-service"
api_key_env = "DATABRICKS_TOKEN"
base_url_env = "FLEET_LLM_BASE_URL"
max_tokens = 16384
timeout_seconds = 90
num_retries = 3
temperature = 0
cache = false
```

Set the two variables in `.env`:

```bash .env theme={null}
DATABRICKS_TOKEN=<your-databricks-token>
FLEET_LLM_BASE_URL=https://<workspace-host>/ai-gateway/mlflow/v1
```

The Databricks DeepSeek endpoint requires an `https` base URL whose path is exactly `/ai-gateway/mlflow/v1`.

## Use a different provider

<Steps>
  <Step title="Set the model ID">
    Set `model` in each role. Fleet adds the `openai/` prefix when the ID has no provider prefix, so `gpt-4o-mini` becomes `openai/gpt-4o-mini`.
  </Step>

  <Step title="Name the credential variable">
    Set `api_key_env` to the name of the environment variable that holds the key. Fleet never reads secrets directly from `config/fleet.toml`.
  </Step>

  <Step title="Set the endpoint">
    Set either `base_url` to a literal URL or `base_url_env` to a variable name. You can't set both in one role.
  </Step>

  <Step title="Restart Fleet">
    Fleet reads `config/fleet.toml` only at startup.
  </Step>
</Steps>

```toml config/fleet.toml theme={null}
[llm.root]
model = "openai/gpt-4o"
api_key_env = "OPENAI_API_KEY"
base_url = "https://api.openai.com/v1"

[llm.sub]
model = "openai/gpt-4o-mini"
api_key_env = "OPENAI_API_KEY"
base_url = "https://api.openai.com/v1"
temperature = 0
```

## Role settings

| Key | Description |
| - | - |
| `model` | Model ID sent to the provider. |
| `api_key_env` | Environment variable that holds the API key. |
| `base_url`, `base_url_env` | Endpoint URL, or the variable that holds it. |
| `max_tokens` | Maximum output tokens. Must be at least 1. |
| `timeout_seconds` | Request timeout. Defaults to 300 for Root and 90 for Sub. |
| `temperature` | Sampling temperature. |
| `reasoning_effort` | `none`, `low`, `medium`, or `high`, for models that support it. |
| `num_retries` | Provider retries. Defaults to 3. |
| `cache` | Enables the DSPy response cache. |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.